llama.cpp b11412 Fixes Computation Graph Reallocation Issues in K-Pool Models
llama.cpp release b11412 resolves an issue where k-pool architecture models, including Qwen4exp and GLM5-next, triggered unexpected runtime graph reallocations during inference. The fix enforces consistent computation graph topologies across execution states by removing dynamic branching based on sequence-sharing conditions. Unexpected graph reallocations degrade inference performance and cause runtime crashes under strict memory allocation checks like GGML_SCHED_DEBUG_REALLOC=1. Ensuring static graph shapes allows llama.cpp to maintain high performance and stable memory consumption during batched multi-sequence processing. The patch eliminates reliance on the dynamic get_kpool_cache_safe graph API and updates the models to always use static bounds for scatter and gather operations. Additionally, it fixes a CUDA MoE weighted reduction bug where empty ubatches dropped allocation dependencies and caused node count mismatches.
## BACKGROUND
llama.cpp relies on the GGML memory allocator, which pre-allocates memory for the full execution context based on a measured worst-case computation graph to avoid allocation overhead during hot decoding paths. If the graph topology changes dynamically at runtime due to varying token counts or sequence states, the scheduler is forced to reallocate memory mid-execution.