OpenAI Repeatedly Modifies GPT-6 Astra Benchmark Data Following Launch Post
OpenAI repeatedly altered the benchmark evaluation scores published in its GPT-6 Astra announcement blog post shortly after its initial release on September 3. These post-launch adjustments affected Astra's reported hallucination rates, coding scores, and even the benchmark results of competitor models like Anthropic's Fable 5.1. The frequent shifting of scores highlights ongoing concerns about the reliability, transparency, and consistency of LLM evaluation benchmarks across the AI industry. It underscores how vendors can engage in 'benchmaxxing' by altering test setups, tool availability, or checkpoints to market superior performance figures. Astra's hallucination rate initially dropped from 4.2% to 2% before reverting back to 4.2%, while its ARC-AGI-3 score reached 99.99% under an enhanced tool framework compared to 63% under standard conditions. OpenAI acknowledged that published scores represent maximum obtainable performance under optimal compute conditions rather than standard user experience.
## BACKGROUND
Large language model (LLM) evaluation benchmarks are standardized test suites used by researchers and developers to evaluate and compare AI models across domain tasks such as reasoning, mathematics, and coding. However, benchmark results can vary significantly depending on prompt design, system tool availability, inference settings, and test execution environments. Furthermore, issues like data contamination—where benchmark test data leaks into model training sets—and parameter tuning to maximize leaderboard rankings complicate fair evaluations.