llama.cpp Release b11413 Optimizes Vulkan Backend for Sparse Flash Attention
llama.cpp release b11413 introduces Vulkan backend optimizations specifically for sparse flash attention using quantized Key-Value (K/V) caches. The update refactors the index compaction process to use a single-scan segment approach, eliminating execution bottlenecks during decoding. Running LLMs on cross-platform consumer hardware via Vulkan often faces memory bandwidth and compute bottlenecks during long-context generation. These optimizations reduce overhead for quantized K/V caches, facilitating faster and lower-memory inference across AMD, Intel, and mobile GPUs. Previously, Vulkan sparse index compaction ran one workgroup per mask row with iterative barriers, which at 128k sequence cells cost more execution time than the attention pass itself. The new implementation splits rows into contiguous segments per subgroup with ballot counting over coalesced loads, relying on a single scan pass to calculate output offsets while maintaining ascending order.
## BACKGROUND
llama.cpp is an open-source inference engine designed for running Large Language Models efficiently on varied local hardware. Quantizing the K/V cache reduces VRAM footprint by storing attention keys and values at lower precision (such as INT8 or INT4), while sparse flash attention avoids redundant calculations across long context sequences to speed up generation.