llama.cpp Release b11475 Fixes Metal Backend Operator Fusion Bug
llama.cpp release b11475 fixes a bug in the Metal backend operator fusion logic where combining matrix multiplication (MUL_MAT) and addition (ADD) failed when both addition operands were matrix multiplications. The update aligns the compute encoder logic with operand identity selection and adds new test coverage in test-backend-ops.cpp. This bug fix resolves silent output accuracy degradation on Apple Silicon GPUs for specific architectures, such as Clef decision models, where output probabilities previously collapsed toward uniform distribution. Apple device users relying on Metal acceleration will now receive correct outputs matching the CPU reference backend without needing to disable fusion optimizations. When evaluating expressions like x = W1 @ u + W2 @ v, the Metal kernel misidentified the residual buffer and read unwritten memory, causing 27 out of 28 unit test cases to fail before this fix. While setting GGML_METAL_FUSION_DISABLE=1 served as a temporary workaround, release b11475 resolves the issue natively within the Metal compute kernel encoder.
## BACKGROUND
Operator fusion is a machine learning compiler optimization that merges multiple consecutive operations into a single GPU kernel execution to reduce memory traffic and latency. The GGML framework provides a Metal backend to run hardware-accelerated tensor operations on Apple devices.