SWE-sweep: A New Benchmark for Proactive AI Bug Finding and Fixing
Researchers from Meta, Stanford, Harvard, and UW have released SWE-sweep, an open-source benchmark designed to evaluate how well AI agents autonomously discover and patch hidden real-world bugs across entire codebases. Unlike traditional reactive debugging benchmarks, SWE-sweep tasks models with proactively searching repository code to fix issues before users run into them. This benchmark shifts the evaluation of coding AI agents from reactive bug-fixing to proactive codebase maintenance. It reveals that current state-of-the-art models still score surprisingly low when asked to discover hidden software defects without explicit user reports. Models are evaluated against a hidden, filtered set of real-world software bugs that are verified to be discoverable directly from the codebase. All code and benchmark assets are open-sourced under the MIT license, with GPT-5.6 Luna (xhigh) currently highlighted for its cost efficiency on the leaderboard.
## BACKGROUND
Standard AI coding benchmarks typically provide language models with a detailed issue description or bug report and measure their ability to generate a patch. However, real-world software engineering requires proactive static analysis, edge-case auditing, and repository-wide exploration to prevent bugs before deployment.