llama.cpp Release b11480 Adds GPU Caching for Host-Stored MoE Experts
llama.cpp release b11480 introduces a GPU caching mechanism for Mixture of Experts (MoE) model parameters retained in CPU host memory. The update introduces `llama_moe_cache_ptr` to optimize expert retrieval during inference. Running large MoE models often requires offloading expert layers to CPU system RAM due to GPU VRAM limits. By caching frequently activated host experts directly on the GPU, llama.cpp reduces data transfer bottlenecks over the PCIe bus and improves text generation speed. The optimization was integrated via PR #29887, adding dynamic pointer management for host-resident MoE experts. Pre-built release binaries were published for a wide range of platforms, including macOS, Linux (CUDA, Vulkan, ROCm, SYCL), Windows, Android, and Snapdragon hardware.
## BACKGROUND
Mixture of Experts (MoE) is an LLM architecture where only a subset of specialized sub-networks, called experts, are routed to process each token. Because only a fraction of total model weights are active per token, offloading inactive experts to CPU memory allows consumers to execute massive models that would otherwise not fit into VRAM.