llama.cpp Release b11313 Re-Enables Split-Tensor Mode for Qwen Models
llama.cpp release b11313 re-enables split-tensor mode (`-sm tensor`) for Qwen model architectures by correcting backend device assignment during computation graph splitting. The update fixes a backend assertion crash that previously occurred when handling host-resident embeddings and memory tensor reshapes. Tensor splitting across multiple GPUs is vital for high-performance multi-GPU inference in llama.cpp, significantly boosting GPU utilization and inference speed. Fixing this scheduling issue allows users running Qwen-based models on multi-GPU systems to utilize tensor-parallel inference without experiencing stability crashes. The issue stemmed from an assertion failure (`GGML_ASSERT(ggml_backend_buffer_is_meta)`) when the graph scheduler failed to assign a `REPEAT` node to the GPU device because it was blocked by CPU-bound gather nodes. Expanding `hc_init` immediately after creation moves the node directly before the device assignment step, matching the implementation strategy used for DeepSeek models.
## BACKGROUND
llama.cpp is a popular open-source framework for running Large Language Models locally across diverse hardware backends, including CPU, CUDA, and Vulkan. Its underlying tensor library, ggml, uses a graph scheduler to partition workload graphs between host CPUs and dedicated accelerators or across multiple GPUs in split-tensor mode.