The trick: Moved Ruler
GPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week.
Nothing changed except whose chart it was.
the week's coding-agent launches stack neatly on DeepSWE v1.1: GPT-5.6 Terra 69.6, Grok 4.6 High 65.9, Gemini 3.7 Flash 65.3, Muse Spark 1.2 59.3. Aggregators and social posts quote these as one ranking.
Before you read on. Your call?
TRUE, BUT
One name, many rulers
no shared harness produced those numbers. Google's launch chart, Meta's launch chart, and xAI's launch table each ran a same-named suite independently, and the same models land in both of the two charts that overlap. That overlap is the tell. GPT-5.6 Terra: 69.6 on Google's chart, 64.8 on Meta's. Muse Spark 1.2: 59.3 on Meta's chart, 54.9 on Google's. Roughly five points of daylight per model, and in each case the chart owner's rival scores lower on the owner's chart. As the one careful comparison in circulation puts it, cross-chart readings are 'suggestive but not proof, since it's two different labs running the same-named suite independently', or shorter: 'a shared benchmark name is not a shared benchmark'. The individual numbers are ordinary first-party launch stats. The leaderboard assembled from them is fiction with a spreadsheet aesthetic.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →
Receipts
- Supports blog.google:
FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%)
- Supports orcarouter.ai:
Claude Opus 5 65.0%, GPT-5.6 Terra 64.8%, Muse Spark 1.2 59.3%
- Context explainx.ai:
GPT-5.6 Terra 69.6% Gemini 3.7 Flash 65.3% Muse Spark 1.2 54.9%
- Refutes explainx.ai:
suggestive but not proof, since it's two different labs running the same-named suite independently
- Refutes explainx.ai:
a shared benchmark name is not a shared benchmark
- Context explainx.ai:
SpaceXAI's own Grok 4.6 launch table separately reported Grok 4.6 High at 65.9% and GPT-5.6 Sol Max at 73% on the same-named DeepSWE V1.1
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.