llama.cpp Release b11440 Fixes Scheduler Re-allocation Bug in Speculative Decoding
llama.cpp release b11440 fixes a graph scheduler re-allocation issue that occurred during speculative decoding initialization when NextN extraction flags changed. The update invalidates the reserved memory schedule when these flags change, forcing a clean re-reservation matching the new compute graph shape. This fix prevents runtime assertion failures (specifically GGML_SCHED_DEBUG_REALLOC) when running speculative decoding with Multi-Token Prediction (MTP) enabled. It ensures memory stability and correct execution for advanced LLM inference configurations. The issue stemmed from unmasked NextN extraction retaining every token through the final layer rather than cropping to output rows, causing the first decode step to reallocate memory dynamically. When subsequent wider batches were processed, it violated scheduler reservation assertions, which is now resolved by explicitly triggering a compute schedule re-reservation.
## BACKGROUND
llama.cpp is a popular open-source C/C++ framework for running LLMs locally across various hardware backends. Speculative decoding speeds up LLM inference by generating draft tokens rapidly and verifying them in batch, with Multi-Token Prediction (MTP) allowing models to draft multiple future tokens using native architectural heads. The underlying GGML library uses a memory scheduler to pre-allocate memory for compute graphs to maximize performance, requiring precise tracking of tensor shapes across compute passes.