OpenAI Launches GPT 6.1 Sol with Reduced API Caching Costs and Tier Updates
OpenAI has introduced GPT 6.1 Sol in Codex and ChatGPT Work, delivering performance close to its high-end Astra model at a lower operational cost. Alongside the model release, OpenAI updated its subscription tier limits and reduced cached input pricing by 50% down to $0.10 per million tokens. Drastic cuts to prompt caching prices significantly lower the cost of running repetitive workflows and long-context applications in tools like Codex. However, concurrent plan changes and price cuts highlight how token economics and margin competition are becoming the primary battleground for major AI providers. Cached input pricing drops to $0.10 per million tokens, representing a 50% discount compared to GPT-6 Sol cached pricing and a 95% reduction from standard input rates. On subscriptions, OpenAI introduced a $500/month Pro 500 plan while reducing the included usage multiplier for $200/month Pro users in Codex and Work from 20x down to 10x.
## BACKGROUND
LLM prompt caching is an optimization technique that stores previously processed input tokens so that subsequent API requests sharing the same prefix do not need to be recomputed. By reusing cached tokens, AI providers enable developers to run long-context applications with significantly lower latency and reduced input costs.