llama.cpp Release b11361 Adds SystemOne API and New Model Architectures
llama.cpp release b11361 introduces a new /v1/systemone server API endpoint along with support for several System One model architectures including laya, julia-1, lev, openjev, and kev. The update adds conversion scripts for these models, vision support, and shared prompt prefix capability. This update expands llama.cpp beyond traditional conversational language models to support structured, fast-decision System One AI architectures. Developers running local inference infrastructure can now serve these decision-making models directly through llama.cpp's lightweight HTTP server. PR #29818 includes a tiny openjev model for testing, updated developer documentation, and clarifies that date_facts is currently not supported. Pre-built binaries are provided for a wide range of backends, including macOS, Linux (CUDA 12/13, Vulkan, SYCL, ROCm, Snapdragon), and Android.
## BACKGROUND
llama.cpp is a popular open-source C/C++ framework for running Large Language Models locally with high efficiency across diverse hardware. System One models, such as TypeSafe AI's Jev series, are specialized models optimized to evaluate state and return fast, strongly-typed decisions directly usable by software, contrasting with traditional generative language models.