llama.cpp Releases Build b11304 with Fix for Training KV Cache Handling
llama.cpp has released automated build b11304, introducing a fix to properly handle Key-Value (KV) cache states during model training (PR #28520). The release provides pre-compiled binaries across a wide array of operating systems and hardware backends, including macOS, Linux, Windows, and Android. Ensuring correct KV cache management during training prevents state corruption and improves numerical consistency when fine-tuning models directly within llama.cpp. Although build b11304 is a routine release, it enhances reliability for developers using llama.cpp's built-in training utilities. The update includes pull request #28520 alongside minor follow-up tweaks for KV handling. Pre-built binaries cover diverse accelerator platforms, including CUDA 12/13, Vulkan, ROCm 10.0, OpenVINO, SYCL, and Snapdragon (CPU, Adreno GPU, and Hexagon NPU).
## BACKGROUND
llama.cpp is an open-source C/C++ execution engine designed for lightweight, highly efficient inference and training of Large Language Models (LLMs) on consumer and server hardware. The Key-Value (KV) cache stores calculated attention key and value vectors across token positions to avoid repeating computationally expensive attention matrix operations.