llama.cpp Release b11418 Adds Vision Input Support for Clef Model
Open-source LLM inference framework llama.cpp released build b11418, introducing vision input capabilities for the Clef decision model in its server component. This update extends previous text-only support by allowing llama-server to process visual data alongside text inputs. Adding multimodal vision capabilities to Clef allows developers to run automated decision-making models locally using visual contexts like screenshots. This expands the scope of lightweight, privacy-focused local AI deployments without needing cloud services. The release includes pull request #29969, which refactors internal batching structures and token positioning to handle multi-dimensional embeddings and image tokens efficiently. Pre-compiled binary packages across platforms including macOS, Linux, Windows, Android, and Snapdragon were updated as part of the release.
## BACKGROUND
llama.cpp is a high-performance C/C++ engine designed for running LLMs locally across diverse hardware, including GPUs, CPUs, and mobile platforms. Clef is a decision model architecture designed to evaluate structured state inputs—such as text or UI screenshots—and compute probabilities for specific actions in a single forward pass.