~/LLAMA CPP/llama-cpp-release-b11365-fixes-cpu-memory-aliasing-bug-in-softmax-backward

llama.cpp Release b11365 Fixes CPU Memory Aliasing Bug in Softmax Backward Operations

llama.cpp has released patch version b11365, which resolves an issue in the CPU backend where the `soft_max_back` operation produced incorrect output when destination memory aliased the input source memory. The update replaces a multi-step execution sequence with a fused loop that reads all inputs before modifying memory. This bug fix prevents silent numerical corruption during backward pass operations on CPU architectures when graph memory optimizations are enabled. Ensuring correct gradient calculations is critical for model training, fine-tuning, and backpropagation tasks built on top of the GGML ecosystem. The error occurred because `GGML_OP_SOFT_MAX_BACK` allowed in-place memory allocation, causing the graph allocator to map `dst` to `src1`, which led early computation steps to overwrite values needed by later steps. CUDA was unaffected because its reduction kernel completes reads before writes, and a regression test was added to verify correct CPU memory aliasing behavior.

## BACKGROUND

llama.cpp is an open-source C/C++ machine learning library built on top of the GGML tensor framework to enable efficient LLM processing across commodity hardware. To minimize RAM footprint, computation graph allocators frequently reuse memory buffers (aliasing) for in-place operations where output tensors share memory space with input tensors.

## REFERENCES

## KEYWORDS

#llama.cpp#ggml#ai-infrastructure#bug-fix

$ subscribe --daily

llama.cpp Release b11365 Fixes CPU Memory Aliasing Bug in Softmax Backward Operations | Daily News