Subscribe

The trick: Moved Ruler

The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified.

Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.

Issue 1319 August 20266 receipts4 min

coding is nearly a solved problem for frontier models. The number carrying that story is SWE-bench Verified, where the top of the board is now packed near 96%, Claude Opus 5 in front at 96%, a figure that reads as almost every real bug fixed.

Before you read on. Your call?

that benchmark is worn out and leaky. benchlm's own leaderboard note says the score is 'nearing saturation for frontier models,' meaning it can no longer tell the best models apart, and OpenAI publicly stopped using SWE-bench Verified in February 2026 after an audit found, in the auditors' words, that '59.4% of audited problems contain flawed test cases that reject correct solutions.' The honest successor is SWE-bench Pro, which, unlike Verified, 'uses actively maintained repositories with no public ground-truth leakage.' On Pro the ceiling falls to 80.3%, and the spread that Verified had flattened reopens: GPT-5.6 Sol, near the top of every coding conversation, sits at 64.6%. The models are strong. Solved is a word the quoted benchmark can no longer support.

96%CLAUDE OPUS 5 ON SWE-BENCH VERIFIED
80.3%THE TOP SCORE ON CONTAMINATION-RESISTANT SWE-BENCH PRO
64.6%GPT-5.6 SOL ON SWE-BENCH PRO
59.4%SHARE OF THE VERIFIED TASKS OPENAI AUDITED THAT HAD FLAWED TESTS

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“The 96% SWE-bench score labs quote is SWE-bench Verified, which is saturated and which OpenAI abandoned after finding 59.4% of audited tasks had flawed tests. On the contamination-resistant SWE-bench Pro the top model scores 80.3% and GPT-5.6 Sol scores 64.6%. Coding is not solved; the ruler is just leaky.”

Receipts

  1. Supports benchlm.ai: Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's August 2026 update with 96%
  2. Context benchlm.ai: suggesting this benchmark is nearing saturation for frontier models
  3. Refutes benchlm.ai: Claude Mythos 5 leads the SWE-bench Pro leaderboard on BenchLM's August 2026 update with 80.3%
  4. Context codingfleet.com: Pro uses actively maintained repositories with no public ground-truth leakage
  5. Context codingfleet.com: above GPT-5.6 Sol (64.6%), at $2/$6 per 1M tokens
  6. Context web.archive.org: First, 59.4% of audited problems contain flawed test cases that reject correct solutions.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.