~/LLAMA CPP/llama-cpp-release-b11474-adds-glm5-next-multi-token-prediction-and-optimizes

llama.cpp Release b11474 Adds GLM5-Next Multi-Token Prediction and Optimizes Inference

llama.cpp release b11474 introduces Multi-Token Prediction (MTP) graph support for GLM5-Next models and optimizes headless forward passes by cropping execution graphs. It also enables loading split MTP-only and trunk-only GGUF files to run draft models efficiently. Multi-Token Prediction accelerates LLM generation by allowing models to predict multiple tokens per inference step or drive speculative decoding. By avoiding unnecessary compute on empty output rows during draft context prefill, execution latency for 4-token catch-up passes dropped from 6.9 ms to 2.9 ms. The update crops attention output and block inputs before the position-wise FFN and shared output head when no output rows are requested. It also resolves edge cases in partial recurrent rollbacks and ensures correct hidden-row extractions during multi-token generation.

## BACKGROUND

Multi-Token Prediction (MTP) is an architectural approach where a language model uses auxiliary modules to predict multiple future tokens at once rather than just the immediate next token. This design is widely used in speculative decoding, where secondary heads quickly draft candidate tokens for the primary model to verify.

## REFERENCES

## KEYWORDS

#llama-cpp#llm-inference#performance-optimization#open-source-ai

$ subscribe --daily

llama.cpp Release b11474 Adds GLM5-Next Multi-Token Prediction and Optimizes Inference | Daily News