~/LLAMA CPP/llama-cpp-adds-gpu-cache-for-moe-experts-stored-in-host-memory

llama.cpp Adds GPU Cache for MoE Experts Stored in Host Memory

A new pull request (#29887) in llama.cpp introduces a GPU caching mechanism for Mixture of Experts (MoE) models whose expert weights are offloaded to host RAM. This allows frequently activated experts to remain in VRAM, significantly speeding up inference on systems with limited GPU memory. Running large MoE models traditionally requires substantial VRAM or suffers severe performance drops when offloading experts to CPU memory due to PCIe bandwidth bottlenecks. By caching active experts in VRAM, local AI users can achieve vastly improved token generation speeds on consumer hardware. The optimization dynamically manages GPU VRAM to keep recently or frequently used expert sub-networks resident on the graphics card, avoiding redundant transfers over the host-to-device bus. This strategy mitigates latency bottlenecks during multi-token generation when consecutive tokens route to overlapping experts.

## BACKGROUND

Mixture of Experts (MoE) is an LLM architecture that replaces heavy feed-forward layers with multiple smaller 'expert' sub-networks, routing each token to only a subset of experts. While MoE models offer high efficiency, their total parameter count requires substantial memory, often forcing users to offload inactive experts to slower system RAM.

## REFERENCES

## KEYWORDS

#llama.cpp#MoE#LLM Inference#GPU Cache#Open Source AI

$ subscribe --daily

llama.cpp Adds GPU Cache for MoE Experts Stored in Host Memory | Daily News