OpenAI's GPT-5.6 Terra and Luna Join Agent Arena Leaderboard
OpenAI's GPT-5.6 Terra and Luna (xHigh variants) have joined the Agent Arena leaderboard, ranking 15th and 17th respectively. These rankings are determined based on real-world agentic sessions evaluated by the global community. This benchmark provides a practical comparison of OpenAI's multi-tier GPT-5.6 model stack against other frontier models in real-world agentic workflows. It helps developers understand the performance-to-cost trade-offs of the balanced (Terra) and lightweight (Luna) variants compared to the flagship Sol model. While the flagship GPT-5.6 Sol ranks in the top 2 with a +10.1% net improvement, Terra (+4.0%) and Luna (+3.3%) show more modest gains, placing them near Claude Opus 4.8 (+3.5%). These xHigh variants represent specific configurations optimized for high-performance agentic tasks.
## BACKGROUND
OpenAI released GPT-5.6 as a three-tier model stack consisting of Sol (for demanding workloads), Terra (for everyday professional use), and Luna (a faster, lower-cost option). Agent Arena is a specialized leaderboard that benchmarks AI models based on how much they improve actual productivity in agentic environments rather than just conversational quality.