llama.cpp Release b11481 Adds Support for Cohere 2 Vision Models
Release b11481 of llama.cpp introduces support for Cohere 2 vision models within its multimodal (mtmd) subsystem via pull request #30062. This allows users to execute Cohere 2 vision-language capabilities directly across supported desktop and mobile devices. Adding Cohere 2 vision support expands the range of multimodal open-source models that can run locally on consumer hardware without reliance on cloud APIs. It further reinforces llama.cpp as the primary cross-platform runtime engine for local AI inference. The update includes optimizations such as linear layer fusion (`fused linear_1`) and cleanups to redundant tensor mappings. Binaries are available across macOS, Linux, Windows, and Android, supporting backends like CUDA 12/13, Vulkan, ROCm, SYCL, OpenVINO, and Snapdragon hardware.
## BACKGROUND
llama.cpp is a popular C/C++ open-source inference library designed to run large language models (LLMs) efficiently on consumer-grade hardware. Its multimodal module (`mtmd`) extends this functionality by enabling vision encoders to process images alongside text prompts.