llama.cpp Release b11388 Adds Activation Statistics for GGUF Importance Matrices
llama.cpp released build b11388, introducing activation-based statistics calculation for GGUF importance matrices (imatrix) alongside refactored draft model processing using llama_batch_ext. This update helps developers evaluate activation behaviors during quantization more precisely, leading to better quality control for compressed GGUF models. It also refines draft model handling to support more efficient speculative decoding workflows in local AI inference. The release adds metrics such as Euclidean–Cosine Score (ECS), two-tailed ZD score, and per-layer L2 norms, accessible via a new --activation-statistics flag designed to prevent doubling default imatrix file sizes. Additionally, external NextN draft model processing (-md / --model-draft) has been updated to leverage llama_batch_ext.
## BACKGROUND
llama.cpp is a widely used open-source framework for running large language models locally on consumer hardware using the GGUF file format. To optimize model size while maintaining response quality, it employs quantization techniques guided by importance matrices (imatrix) that calibrate weight importance. It also supports speculative decoding, which uses smaller draft models to guess upcoming tokens and speed up generation on the main model.