llama.cpp Prepares Support for Upcoming MiniCPM-V 4.7 Vision-Language Model
A pull request (#29416) was submitted to llama.cpp by contributor tc-mb to add support for the upcoming MiniCPM-V 4.7 vision-language model architecture. This PR follows recent sightings of an unreleased MiniCPM-V 4.7 model variant briefly hosted on Hugging Face. Adding early architecture support in llama.cpp ensures local AI enthusiasts and developers can run MiniCPM-V 4.7 as soon as OpenBMB officially releases model weights. MiniCPM-V models are popular in the open-source ecosystem for delivering strong multimodal performance on hardware with limited resources. MiniCPM-V 4.7 pairs a SigLIP vision encoder with a window-attention merger and a Qwen3.5 language model backbone, supporting both 4x and 16x visual downsampling modes. While code implementation for inference is being finalized in llama.cpp, users must wait until OpenBMB officially makes the weights public.
## BACKGROUND
MiniCPM-V is a family of efficient multimodal large language models developed by OpenBMB for fast vision-text inference on edge and desktop hardware. llama.cpp is the de facto standard C/C++ engine for running quantised open-source LLMs locally across CPU and GPU architectures.