llama.cpp Release b11514 Fixes Musa Fast Walsh-Hadamard Transform Bug
llama.cpp has released build b11514, which delivers a targeted bug fix for the Fast Walsh-Hadamard Transform (FWHT) implementation on MUSA backends via PR #30167. This update ensures proper computation on Moore Threads hardware platforms without modifying core APIs. This routine maintenance release ensures stability and mathematical correctness when executing llama.cpp models on domestic Chinese Moore Threads GPUs. It underscores llama.cpp's commitment to supporting a wide array of specialized and alternative GPU architectures worldwide. The release focuses solely on resolving issue #30167 within the MUSA FWHT kernel codebase. Updated pre-built binaries have been generated across standard distribution targets, including Linux, Windows, macOS, Android, and Snapdragon devices.
## BACKGROUND
llama.cpp is a popular open-source C/C++ inference framework optimized for running LLMs on consumer and enterprise hardware. MUSA is a GPU computing platform and software framework developed by Moore Threads as an alternative to NVIDIA's CUDA. The Fast Walsh-Hadamard Transform (FWHT) is an efficient computational algorithm used in mathematical operations and certain model compression or quantization techniques within deep learning.