llama.cpp Adds `/v1/systemone` API Endpoint for System One Models
Pull request #29818 for llama.cpp introduces a new `/v1/systemone` API endpoint to llama-server. This update adds native server support for several open-source System One GGUF models, including OpenJev, Lev, Julia-1, Laya, and Kev-4B. Adding native API support for System One decision models allows developers to run low-latency, typed decision services locally via llama.cpp without relying on proprietary cloud APIs. It expands llama.cpp beyond general text generation into specialized task routing and structured decision-making. The `/v1/systemone` endpoint implements a structured decision API compatible with Jev-style typed questions and output probabilities. Quantized GGUF weights for the supported models are hosted on Hugging Face under the official ggml-org organization.
## BACKGROUND
System One AI models are fast decision-making systems inspired by human intuitive thinking, optimized to produce immediate typed decisions rather than long conversational text. TypeSafe AI popularized this concept with its Jev decision API, prompting open-source efforts like OpenJev to create compatible open-weights alternatives.