llama.cpp b11316 Released with Metal Backend Optimizations for MXFP4
llama.cpp release b11316 updates the Apple Metal backend to perform bfloat16 (bf16) math operations during MXFP4 matrix multiplication. This change optimizes performance and precision during low-bit quantized model inference on Apple hardware. As 4-bit microscaling quantization formats like MXFP4 gain adoption for large language models, optimizing their compute kernels for Apple Silicon is key to fast local inference. Switching to bfloat16 math helps preserve numerical dynamic range while maintaining compute efficiency on modern Apple GPUs. The update incorporates pull request #29770, modifying the Metal compute shaders specifically for MXFP4 matrix multiplication (`mul-mat`). Pre-built binaries were generated across Linux, Windows, macOS, Android, and Snapdragon platforms, though the KleidiAI-enabled macOS arm64 build remains disabled.
## BACKGROUND
llama.cpp is a popular open-source C/C++ framework designed for high-performance LLM inference on edge hardware and personal computers. MXFP4 (Microscaling 4-bit Floating Point) is a low-bit quantization format that compresses neural network weights while preserving accuracy through block-level scaling factors. Metal is Apple's low-level graphics and compute API used by llama.cpp to hardware-accelerate model execution on Apple Silicon.