~/LLAMA CPP/llama-cpp-release-b11313-re-enables-split-tensor-mode-for-qwen-models

llama.cpp Release b11313 Re-Enables Split-Tensor Mode for Qwen Models

llama.cpp release b11313 re-enables split-tensor mode (`-sm tensor`) for Qwen model architectures by correcting backend device assignment during computation graph splitting. The update fixes a backend assertion crash that previously occurred when handling host-resident embeddings and memory tensor reshapes. Tensor splitting across multiple GPUs is vital for high-performance multi-GPU inference in llama.cpp, significantly boosting GPU utilization and inference speed. Fixing this scheduling issue allows users running Qwen-based models on multi-GPU systems to utilize tensor-parallel inference without experiencing stability crashes. The issue stemmed from an assertion failure (`GGML_ASSERT(ggml_backend_buffer_is_meta)`) when the graph scheduler failed to assign a `REPEAT` node to the GPU device because it was blocked by CPU-bound gather nodes. Expanding `hc_init` immediately after creation moves the node directly before the device assignment step, matching the implementation strategy used for DeepSeek models.

## BACKGROUND

llama.cpp is a popular open-source framework for running Large Language Models locally across diverse hardware backends, including CPU, CUDA, and Vulkan. Its underlying tensor library, ggml, uses a graph scheduler to partition workload graphs between host CPUs and dedicated accelerators or across multiple GPUs in split-tensor mode.

## REFERENCES

## KEYWORDS

#llama-cpp#llm-inference#open-source-ai#cpp

$ subscribe --daily

llama.cpp Release b11313 Re-Enables Split-Tensor Mode for Qwen Models | Daily News