llama.cpp Release b11424 Fixes Vulkan Flash Attention Shared Memory Bug
Open-source LLM inference framework llama.cpp released build b11424, featuring a bug fix for an out-of-bounds shared memory write issue (#29988). The fix specifically addresses stability and correctness in the Vulkan backend implementation of Flash Attention. Out-of-bounds memory writes can cause driver crashes, memory corruption, or silent numerical errors during model inference on GPUs utilizing Vulkan. Resolving this issue ensures safer and more reliable Flash Attention execution for users running large language models via Vulkan hardware acceleration. The release includes pre-built binaries across multiple platforms, including Linux, Windows, macOS, Android, and Snapdragon devices with support for CUDA, ROCm, SYCL, and Vulkan backends. Notably, the macOS KleidiAI integration remains temporarily disabled in this release build.
## BACKGROUND
llama.cpp is a popular open-source C/C++ library designed for efficient local inference of Large Language Models (LLMs) across diverse hardware backends. Flash Attention is a memory-efficient attention algorithm that significantly accelerates transformer model inference and reduces memory overhead by tiling operations. Vulkan is a cross-platform graphics and compute API that allows llama.cpp to execute GPU acceleration on non-NVIDIA or integrated hardware graphics processors.