llama.cpp v0.6.0 Released with MTP Speculative Decoding Support
Open-source LLM inference engine llama.cpp has officially released version v0.6.0, introducing support for Multi-Token Prediction (MTP) speculative decoding for Qwen4Exp alongside numerous performance improvements. As one of the most popular local LLM engines, llama.cpp adding native MTP support enables significantly faster inference speeds without requiring a separate draft model. This improves the efficiency of running advanced LLMs on consumer hardware and local deployments. Unlike conventional speculative decoding methods that require an external secondary model to generate candidate tokens, MTP uses the target model's native multi-token prediction head to draft candidate tokens ahead, verifying them in a single forward pass.
## BACKGROUND
llama.cpp is a high-performance C/C++ framework for running LLMs locally across diverse hardware architectures including CPUs and GPUs. Speculative decoding is an acceleration technique that reduces generation latency by predicting multiple tokens per step and verifying them in parallel.