llama.cpp b11254 Fixes Pipeline Parallelism Memory Allocation Bug
llama.cpp release b11254 updates the GGML backend to collect all input tensors after graph splitting rather than during it. This maintains a consistent computation graph composition when processing varying batch types during pipeline parallelism. This fix prevents false backend change detections and unnecessary memory re-reservations that previously caused crashes during pipeline-parallel execution. It enhances stability when running multimodal models that switch between text and vision processing batches across multiple GPUs or devices. Previously, switching batch types (such as token-only versus image batches) shifted leaf nodes in graph_copy, triggering spurious scheduler backend ID changes and improper re-allocations. By collecting inputs from all graph leaf nodes post-split, input tensor composition depends only on which inputs exist rather than which ones are actively consumed.
## BACKGROUND
llama.cpp relies on GGML, an open-source tensor library designed for efficient machine learning inference on commodity hardware. Pipeline parallelism is an optimization technique that splits a large language model's layers across multiple compute devices or execution stages to distribute memory demands.