Subscribe

The trick: Cherry-Picked Slice

The '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.

OpenAI really did retire its own benchmark and the contamination evidence is damning. But 59.4% is the flaw rate of a hand-picked failure pile. As a share of the full benchmark, the confirmed broken tasks are 82 out of 500.

Issue 814 August 20263 receipts4 min

OpenAI retired SWE-bench Verified after its audit found 59.4% of tasks had flawed test cases, and every frontier model trained on the solutions

Before you read on. Your call?

the retirement is real, dated February 23, 2026, and the contamination findings are the strongest part. But the 59.4% comes from an audit of 138 problems selected precisely because OpenAI's o3 failed them across 64 runs. Broken tests are unsolvable, so they pile up in exactly that failure set. As a share of the whole 500-problem benchmark, the confirmed flawed tasks are 82, or 16.4%. The right reading is that the top of the benchmark was phantom headroom, not that the whole thing was always garbage.

59.4%SHARE OF THE 138 AUDITED PROBLEMS
27.6%SHARE OF THE DATASET AUDITED
31'ALMOST IMPOSSIBLE' TASKS GPT-5.2 SOLVED ANYWAY

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“'59.4% of which tasks?' The audit only looked at the 138 problems o3 kept failing. Flawed tests live in the failure pile by definition. The confirmed count is 82 of 500. The other 418 were never audited, so the honest phrase is not shown to be broken, which is not the same as proven clean.”

Receipts

  1. Supports web.archive.org: First, 59.4% of audited problems contain flawed test cases that reject correct solutions.
  2. Refutes web.archive.org: We conducted an audit of 138 SWE-bench Verified problems that OpenAI o3 did not consistently solve over 64 independent runs.
  3. Context web.archive.org: In 59.4% of the 138 tasks, the test design or the problem statement itself was flawed.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.