The trick: Cherry-Picked Slice
The '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.
OpenAI really did retire its own benchmark and the contamination evidence is damning. But 59.4% is the flaw rate of a hand-picked failure pile. As a share of the full benchmark, the confirmed broken tasks are 82 out of 500.
OpenAI retired SWE-bench Verified after its audit found 59.4% of tasks had flawed test cases, and every frontier model trained on the solutions
Before you read on. Your call?
TRUE, BUT
27.6%
the retirement is real, dated February 23, 2026, and the contamination findings are the strongest part. But the 59.4% comes from an audit of 138 problems selected precisely because OpenAI's o3 failed them across 64 runs. Broken tests are unsolvable, so they pile up in exactly that failure set. As a share of the whole 500-problem benchmark, the confirmed flawed tasks are 82, or 16.4%. The right reading is that the top of the benchmark was phantom headroom, not that the whole thing was always garbage.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →
Receipts
- Supports web.archive.org:
First, 59.4% of audited problems contain flawed test cases that reject correct solutions.
- Refutes web.archive.org:
We conducted an audit of 138 SWE-bench Verified problems that OpenAI o3 did not consistently solve over 64 independent runs.
- Context web.archive.org:
In 59.4% of the 138 tasks, the test design or the problem statement itself was flawed.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.