llama.cpp Release b11379 Fixes Server Abort Issue
The llama.cpp project released build b11379, which resolves a server component crash by capping the batch size parameter (n_batch) to not exceed the micro-batch size (n_ubatch). This patch prevents unexpected server crashes during LLM inference serving, ensuring greater system stability for applications relying on the llama.cpp HTTP server component. The fix was introduced in PR #29903 to address issue #29902 with assistance from Claude, and included minor cleanup removing redundant tests and embedding conditions.
## BACKGROUND
llama.cpp is a popular open-source C/C++ framework for running LLM inference locally across diverse hardware platforms. In its architecture, n_batch defines the maximum number of prompt tokens processed in a batch, while n_ubatch sets the physical micro-batch size for compute buffer allocation.