OpenAI Introduces GPT-6 Astra Ultrafast Tier Reaching Up to 300 Tokens per Second
At OpenAI DevDay 2026, OpenAI announced the new Ultrafast service tier for GPT-6 Astra, delivering generation speeds up to 300 tokens per second. This tier offers up to an 8x speedup in Codex and up to 6x faster output for API calls. Historically, developers had to trade off model intelligence for lower latency, relying on smaller models to get near-real-time responses. Ultrafast bridges this gap by combining top-tier intelligence with high throughput, transforming applications in real-time coding, interactive customer service, and event-driven automation. API pricing for short context (<272K tokens) is set at $60 per million input tokens and $300 per million output tokens, while long-context pricing doubles. The tier is accessible via API and available to Pro 500 subscribers ($500/month) and ChatGPT Work users, with OpenAI planning a future GPT-6.1 Sol Ultrafast release.
## BACKGROUND
Large language models generate text sequentially token by token, where generation throughput—measured in tokens per second—directly determines responsiveness in interactive systems. OpenAI Codex is an AI agent system fine-tuned for software engineering tasks like code generation and debugging.