llama.cpp Release b11524 Drops mp_21 Target for MUSA Backend
Open-source project llama.cpp has released version b11524, which removes `mp_21` from the default target GPU architectures for the MUSA backend. This modification was submitted by Xiaodong Ye from Moore Threads in pull request #30203. Removing legacy or unsupported compute architectures streamlines compilation and prevents build errors for developers running llama.cpp on Moore Threads GPU platforms. It highlights ongoing maintenance to ensure modern LLM inference tools properly support specialized GPU accelerators. Alongside the MUSA build configuration tweak, release b11524 provides updated binaries across Linux, Windows, macOS, Android, and Snapdragon devices. Additionally, KleidiAI support for macOS Apple Silicon builds remains disabled in this release.
## BACKGROUND
llama.cpp is a widely used open-source C/C++ library designed for efficient inference of Large Language Models (LLMs) across various hardware platforms. MUSA (Moore Threads Unified System Architecture) is a proprietary GPU hardware and software architecture developed by Moore Threads for graphics and AI workload acceleration.