Source: OpenAI Research — 2026-07-08
Summary
OpenAI published an audit of SWE-bench Verified, the widely-used coding-agent benchmark, finding that 27–34% of tasks are broken in some way — overly strict test cases, underspecified prompts, or data contamination. The piece argues that as currently constituted, the benchmark no longer gives a reliable read on real agentic coding capability, and calls for the field to treat headline SWE-bench scores with more skepticism.
Key Takeaways
- 27–34% of SWE-bench Verified tasks were found to have some form of defect: over-strict tests, ambiguous specs, or contamination.
- The audit's conclusion directly challenges how much weight the industry should put on SWE-bench leaderboard rankings for comparing coding agents.
- Likely to reshape how coding-agent benchmark results get reported and scrutinized going forward, given OpenAI's own stake in those leaderboards.
- A methodological, self-critical piece rather than a capability announcement — notable for that alone.