llama.cpp Releases Version 0.6.0 (Build b11429) Across Multiple Platforms
The open-source local LLM inference engine llama.cpp has officially bumped its version to 0.6.0 (build b11429). This release provides updated prebuilt binaries across macOS, iOS, Linux, Android, and Windows for a wide array of CPU, GPU, and NPU execution backends. As llama.cpp is a foundational tool for local Large Language Model (LLM) execution, milestone releases simplify deployment across edge devices and servers. The release expands precompiled binary support for cutting-edge hardware tooling, including CUDA 13, ROCm 10, Vulkan, and Qualcomm Snapdragon NPUs. The release package features binaries tailored for platforms such as Intel/ARM CPUs, NVIDIA CUDA 12/13, AMD ROCm 10.0, Intel SYCL, OpenVINO, and Snapdragon's Hexagon NPU. Notably, the macOS Apple Silicon build with integrated Arm KleidiAI acceleration remains temporarily disabled.
## BACKGROUND
llama.cpp is a lightweight, open-source C/C++ library designed for efficient local inference of large language models across diverse hardware platforms. Arm KleidiAI is an open-source software library providing optimized routines and micro-kernels to accelerate AI workloads on Arm CPUs.