~/LLM BENCHMAR/developer-reports-suspiciously-high-96-score-for-ox-alpha-on-swe-bench

Developer Reports Suspiciously High 96% Score for Ox Alpha on SWE-bench Verified Mini

A developer reported that the mystery AI model "Ox Alpha" achieved a 96% resolution rate (48 out of 50 tasks) on the SWE-bench Verified Mini dataset using the official mini-swe-agent scaffold. However, the developer expressed strong skepticism about the result, suggesting it may be inflated due to data contamination or the specific subset used. This benchmark highlight underscores the ongoing challenges in evaluating LLMs, particularly the risk of data contamination where models memorize training data from popular repositories like Django and Sphinx. It also raises questions about the capabilities of "Ox Alpha," a stealth model offered for free on OpenRouter. The test was conducted locally using the official SWE-bench Docker harness with the mini-swe-agent v2.4.6 scaffold, where Ox Alpha failed only two Django tasks. The developer noted that the 50-task mini subset only contains Django and Sphinx repositories, which are heavily represented in LLM training data, making memorization highly likely.

## BACKGROUND

SWE-bench Verified is a human-filtered subset of 500 software engineering tasks designed to test AI agents on resolving real GitHub issues. The "mini" version is a smaller 50-task subset, and mini-swe-agent is a lightweight, widely adopted AI agent scaffold designed to run these evaluations efficiently.

## REFERENCES

## KEYWORDS

#LLM Benchmarks#SWE-bench#AI Agents#Open Source LLMs

$ subscribe --daily

Developer Reports Suspiciously High 96% Score for Ox Alpha on SWE-bench Verified Mini | Daily News