llama.cpp b11298 Adds Support for DFlash Conversion and Feature Extraction
llama.cpp build b11298 has been released, introducing model conversion and feature extraction support for DFlash. This update enables llama.cpp users to convert and experiment with DFlash, an emerging block diffusion framework designed to accelerate LLM inference via speculative decoding. The update incorporates pull request #29650 created with contributions from Hugging Face developers, along with updated cross-platform release binaries for Linux, Windows, macOS, Android, and Snapdragon platforms.
## BACKGROUND
llama.cpp is a widely used open-source framework written in C/C++ for executing Large Language Models locally across diverse hardware. DFlash is a lightweight block diffusion model architecture created for speculative decoding, which speeds up text generation by generating parallel draft tokens during LLM inference.