Subscribe

The trick: Moved Ruler

GPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week.

Nothing changed except whose chart it was.

Issue 1016 August 20266 receipts3 min

the week's coding-agent launches stack neatly on DeepSWE v1.1: GPT-5.6 Terra 69.6, Grok 4.6 High 65.9, Gemini 3.7 Flash 65.3, Muse Spark 1.2 59.3. Aggregators and social posts quote these as one ranking.

Before you read on. Your call?

no shared harness produced those numbers. Google's launch chart, Meta's launch chart, and xAI's launch table each ran a same-named suite independently, and the same models land in both of the two charts that overlap. That overlap is the tell. GPT-5.6 Terra: 69.6 on Google's chart, 64.8 on Meta's. Muse Spark 1.2: 59.3 on Meta's chart, 54.9 on Google's. Roughly five points of daylight per model, and in each case the chart owner's rival scores lower on the owner's chart. As the one careful comparison in circulation puts it, cross-chart readings are 'suggestive but not proof, since it's two different labs running the same-named suite independently', or shorter: 'a shared benchmark name is not a shared benchmark'. The individual numbers are ordinary first-party launch stats. The leaderboard assembled from them is fiction with a spreadsheet aesthetic.

69.6%GPT-5.6 TERRA ON GOOGLE'S DEEPSWE CHART
64.8%THE SAME TERRA ON META'S DEEPSWE CHART
59.3%MUSE SPARK PER META. GOOGLE'S CHART SAYS 54.9%

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Terra is 69.6 on Google's chart and 64.8 on Meta's. Same model, same benchmark name, five points apart. Every cross-lab gap smaller than that is noise wearing a ranking.”

Receipts

  1. Supports blog.google: FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%)
  2. Supports orcarouter.ai: Claude Opus 5 65.0%, GPT-5.6 Terra 64.8%, Muse Spark 1.2 59.3%
  3. Context explainx.ai: GPT-5.6 Terra 69.6% Gemini 3.7 Flash 65.3% Muse Spark 1.2 54.9%
  4. Refutes explainx.ai: suggestive but not proof, since it's two different labs running the same-named suite independently
  5. Refutes explainx.ai: a shared benchmark name is not a shared benchmark
  6. Context explainx.ai: SpaceXAI's own Grok 4.6 launch table separately reported Grok 4.6 High at 65.9% and GPT-5.6 Sol Max at 73% on the same-named DeepSWE V1.1

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.