DLoop Introduces Looped Speculative Decoding to Accelerate LLM Inference
Researchers from NAVER AI introduced DLoop, a novel speculative decoding framework that executes multiple adaptive drafting loops before triggering target-model verification. By continuing to draft while the small model remains confident, DLoop significantly reduces unnecessary forward passes on the target model during LLM generation. This method resolves a key bottleneck in speculative decoding where the target model verifies draft tokens too frequently, even when the draft model is highly accurate. Across multiple decoding techniques like EAGLE-3 and Domino, DLoop delivers a 5% to 41% wall-clock speedup while maintaining strictly lossless output quality. To maintain high drafting accuracy across multiple unverified iterations, DLoop uses loop-aware training that exposes the draft model to its own hidden states from previous unverified steps. The technique is compatible with both autoregressive and parallel draft models, including multi-token prediction modules and methods like DFlash and DSpark.
## BACKGROUND
Large Language Model (LLM) inference is often memory-bandwidth bound because autoregressive generation generates one token at a time. Speculative decoding addresses this by using a lightweight draft model to quickly propose a sequence of tokens, which the larger target model then verifies all at once in parallel.