llama.cpp Release b11345 Adds Q2_K and Q3_K Quantization for Qualcomm Hexagon DSP
llama.cpp release b11345 introduces support for Q2_K and Q3_K quantization types tailored for the Qualcomm Hexagon DSP backend. The pull request (#29717) was co-authored by Qualcomm engineer Max Krasnyansky. Enabling low-bit quantization on Qualcomm Hexagon DSP allows mobile and edge devices powered by Snapdragon chips to run LLMs with significantly reduced memory footprints. This improvement enhances on-device performance and power efficiency for edge AI applications. The update includes backend code fixes for consistent memory allocation of `src1_row_size` on Hexagon architectures. Updated pre-built binaries have been made available for Snapdragon platforms across Linux arm64 and Android arm64.
## BACKGROUND
llama.cpp is an open-source C/C++ framework for local inference of large language models with high execution efficiency. Quantization methods like Q2_K and Q3_K compress model weights to 2-bit or 3-bit integer representations to drastically lower RAM usage while maintaining acceptable model quality. The Qualcomm Hexagon DSP is a specialized hardware coprocessor built into Snapdragon SoCs for accelerating low-power multimedia and AI workloads.