The trick: Moved Ruler
The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified.
Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.
coding is nearly a solved problem for frontier models. The number carrying that story is SWE-bench Verified, where the top of the board is now packed near 96%, Claude Opus 5 in front at 96%, a figure that reads as almost every real bug fixed.
Before you read on. Your call?
TRUE, BUT
96%
that benchmark is worn out and leaky. benchlm's own leaderboard note says the score is 'nearing saturation for frontier models,' meaning it can no longer tell the best models apart, and OpenAI publicly stopped using SWE-bench Verified in February 2026 after an audit found, in the auditors' words, that '59.4% of audited problems contain flawed test cases that reject correct solutions.' The honest successor is SWE-bench Pro, which, unlike Verified, 'uses actively maintained repositories with no public ground-truth leakage.' On Pro the ceiling falls to 80.3%, and the spread that Verified had flattened reopens: GPT-5.6 Sol, near the top of every coding conversation, sits at 64.6%. The models are strong. Solved is a word the quoted benchmark can no longer support.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →
Receipts
- Supports benchlm.ai:
Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's August 2026 update with 96%
- Context benchlm.ai:
suggesting this benchmark is nearing saturation for frontier models
- Refutes benchlm.ai:
Claude Mythos 5 leads the SWE-bench Pro leaderboard on BenchLM's August 2026 update with 80.3%
- Context codingfleet.com:
Pro uses actively maintained repositories with no public ground-truth leakage
- Context codingfleet.com:
above GPT-5.6 Sol (64.6%), at $2/$6 per 1M tokens
- Context web.archive.org:
First, 59.4% of audited problems contain flawed test cases that reject correct solutions.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.