llama.cpp Release b11368 Adds Probabilistic and Rejection Sampling for Speculative Decoding
Release b11368 of llama.cpp introduces probabilistic draft sampling and target-side rejection sampling for speculative draft decoding and Multi-Token Prediction (MTP). The update enables draft models to sample candidate tokens probabilistically while allowing the primary target model to verify those tokens via rejection sampling. Speculative decoding accelerates LLM inference by generating multiple tokens ahead of time, but traditional greedy sampling can limit generation quality or diversity. Support for probabilistic and rejection sampling ensures that users can achieve faster inference latency while maintaining proper probability distributions in model outputs. The release adds a new flag to enable probabilistic draft sampling (defaulting to greedy decoding) and automatically falls back to argmax sampling for grammar-constrained requests. Technical fixes include separating the random number generator (RNG) stream between draft and target samplers, renormalizing probability distributions after masking, and moving replay logic to the server.
## BACKGROUND
Speculative decoding is an inference-time optimization for large language models that uses a fast, lightweight draft model or architecture extension to propose multiple tokens, which are then verified in parallel by the larger target model. Multi-Token Prediction (MTP) enhances this paradigm by training or adapting models to predict multiple future sequence tokens simultaneously rather than just one next token at a time.