OpenAI Launches Low-Latency Decisions API for Real-Time AI Routing
OpenAI announced the Decisions API at DevDay 2026, a specialized low-latency service engineered to deliver structured classification and routing results within roughly 150 milliseconds. Powered by OpenAI's lightweight Luna model, the API performs single-step evaluations nearly 10 times faster than conventional API calls. The launch marks a transition from open-ended conversational chat to fast, deterministic decision nodes within production software architecture. It enables developers to integrate rapid AI-powered decision logic directly into customer support routing, content moderation, and multi-agent workflow orchestration. Developers provide context (text or image) along with pre-defined questions and a restricted set of valid candidate answers, receiving back selected options alongside confidence scores. The API directly targets high-speed decision-making tasks, competing with existing machine-native solutions like TypeSafe's Jev model.
## BACKGROUND
Standard Large Language Model (LLM) APIs often incur high latency because they generate unstructured text responses token by token. For mission-critical logic like automated routing or business rules, developers require ultra-fast structured outputs rather than open-ended text generation, giving rise to specialized decision-focused endpoints and models.