llama.cpp b11240 Adds Multimodal Embeddings Support to Server API
llama.cpp build b11240 adds support for OpenAI-style typed content inputs—including vision, audio, and video—to its /v1/embeddings server endpoint. This update enables embedding generation for multimodal models like Qwen3-VL-Embedding directly through the local server interface. This enhancement expands the capabilities of open-source local inference servers, enabling developers to build cross-modal retrieval-augmented generation (RAG) and search applications. By adhering to OpenAI's structured content array format, it ensures seamless compatibility with modern multimodal AI application workflows. The endpoint accepts OpenAI content arrays where text parts are concatenated and media items are decoded and spliced into prompts, while preserving legacy string and token formats. Additionally, Key-Value (KV) prefix reuse is disabled for stateless embedding and reranking tasks to prevent requests from incorrectly sharing cached states.
## BACKGROUND
llama.cpp is a popular open-source framework designed for high-performance LLM inference across diverse CPU and GPU hardware platforms. Embeddings translate unstructured inputs such as text, images, or audio into mathematical vectors, facilitating semantic search, clustering, and retrieval in AI pipelines.