Running Qwen 3.6 35B Vision Model with 131K Context on a 6GB GPU
A local LLM enthusiast demonstrated how to run the Qwen 3.6 35B A3B multimodal vision model with a 131K context window on an entry-level RTX 2060 6GB GPU using llama.cpp. By offloading 39 MoE CPU experts and the vision projector to system RAM, the setup achieved decoding speeds of 15–23 tokens per second while keeping VRAM usage under 5.2 GB. This benchmark proves that large 35-billion parameter Mixture-of-Experts (MoE) vision models can run locally on budget consumer hardware. It highlights how expert and projector offloading strategies democratize access to long-context, multimodal AI without requiring enterprise-grade GPUs. The configuration utilized a Q4_K_M quantized model paired with a Q8_0 KV cache on a host machine equipped with an RTX 2060 6GB, 32GB DDR4 RAM, and an Intel i5-10400F CPU. Running the vision projector on the CPU (`--no-mmproj-offload`) saved approximately 1GB of VRAM to prevent out-of-memory crashes, while prefill speed stabilized around 485 tokens per second even at a 90K context depth.
## BACKGROUND
Mixture-of-Experts (MoE) architectures divide model weights into specialized sub-networks ('experts') and route inputs dynamically, making it possible to offload inactive expert layers to system CPU RAM instead of holding them in GPU memory. Frameworks like llama.cpp leverage model quantization (such as Q4_K_M) and KV cache compression to significantly lower VRAM usage during high-context inference.