llama.cpp Release b11518 Adds 128/96 Flash Attention Kernels for Apple Metal
llama.cpp released version b11518, which introduces 128/96 Flash Attention kernels to its Metal backend for Apple devices. The update also enables the framework to accept every compiled kernel pair for better execution flexibility. This update optimizes memory usage and speeds up LLM inference on Apple Silicon devices by expanding hardware-accelerated Flash Attention kernel support. It allows models using 128 or 96 head dimensions to run more efficiently on macOS and iOS devices. The update comes via pull request #30209 and specifically targets Metal compute kernels for attention head dimensions of 128 and 96. Along with the release, pre-built binaries were made available for macOS, iOS, Linux, Android, and Windows across CPU, CUDA, Vulkan, and Snapdragon backends.
## BACKGROUND
llama.cpp is an open-source framework designed for local, high-performance LLM inference across a wide range of hardware architectures. Metal is Apple's compute framework for accelerating graphics and general GPU tasks on Apple Silicon hardware. Flash Attention is a memory-efficient algorithm that uses tiling to minimize High Bandwidth Memory (HBM) I/O, enabling faster Transformer model computation.