llama.cpp b11450 Release Adds Tensor Split Support to RPC
llama.cpp release b11450 introduces support for tensor split mode (`-sm tensor`) over Remote Procedure Call (RPC), enabling tensor partitioning across network nodes. The release also fixes flush handling for Apple RDMA and includes optimizations to stop dispatcher thread spinning. Extending tensor splitting to RPC allows distributed inference setups to divide individual model layers across multiple networked machines rather than relying solely on layer-by-layer distribution. This expands flexibility and memory efficiency for running large language models on clustered hardware. Because this release bumps the RPC major version, users must update both client and server nodes simultaneously due to breaking protocol changes. Additional backend changes include moving `graph_uids` to `rpc_dispatcher` and cleaning up meta backend logic.
## BACKGROUND
llama.cpp is an open-source C/C++ framework for running large language models locally with high efficiency. Its built-in RPC capabilities allow users to link multiple computers together over a network to combine memory and compute power for hosting large models.