llama.cpp b11387 Fixes N-Gram Draft Rejection Bug in Speculative Decoding
llama.cpp release b11387 fixes a bug where n-gram draft tokens in speculative decoding were incorrectly rejected when the sampling temperature was set above zero after context truncation. The fix was contributed in PR #29924 by NVIDIA engineer Pranesh Gonegandla. N-gram speculative decoding speeds up LLM inference by matching previously generated text patterns without needing a secondary draft model. Fixing this bug ensures that draft acceptance logic functions properly at non-zero temperatures, preventing unnecessary token re-generation and performance penalties. The release includes pre-compiled binaries across Linux, Windows, macOS, Android, and iOS supporting backends like CUDA 12/13, Vulkan, ROCm 10.0, OpenVINO, SYCL, and Snapdragon hardware.
## BACKGROUND
llama.cpp is a popular open-source project enabling high-performance LLM inference across diverse hardware backends. Speculative decoding is an acceleration technique where candidate tokens (drafts) are proposed quickly—such as via n-gram sequence matching—and verified by the main LLM in parallel.