Subscribe

The trick: Announced Not Shipped

Grok 4.6 posts a 1753 and a number one.

One is a real benchmark. The other is a tweet.

Issue 713 August 20269 receipts4 min

Musk says Grok 4.6 reaches 1753 ELO and is number one on Databricks' OfficeQA benchmark, at half the price of rivals.

Before you read on. Your call?

the 1753 is real, and it comes from Artificial Analysis, not just Musk. But it is one benchmark, GDPval-AA v2, where Grok sits behind Claude Opus 5 and inside overlapping confidence intervals with Fable 5 and Qwen. On the overall Intelligence Index it scores 61, tied with GPT-5.6 Sol for third, behind Opus at 63 and Fable at 62.

The twist

the number one Musk tweeted is an OfficeQA Pro V2 run with Databricks' Genie harness, and in one analysis's words that number 'is not in SpaceXAI's published table.' xAI's model card does show a number one on the older OfficeQA Pro v1, 63.2 percent, but xAI ran that itself, and Databricks' own leaderboard does not list Grok 4.6 at all. The part that actually holds is the boring part: Grok 4.6 is cheap, about $0.84 a task, with list prices of $2 in and $6 out per million tokens against $5 and $25 for Opus 5.

1753ELO
61INTELLIGENCE INDEX
$0.84COST PER TASK

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Announced Not Shipped: claimed, never released. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“1753 is real and it is tied for third. The number one is a benchmark run Databricks has not published, and Grok 4.6 is not on Databricks' own leaderboard.”

Receipts

  1. Context artificialanalysis.ai: achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5
  2. Context artificialanalysis.ai: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62)
  3. Refutes explainx.ai: That number is not in SpaceXAI's published table.
  4. Context explainx.ai: Treat it as a founder claim until Databricks or a third party reproduces it.
  5. Supports cryptobriefing.com: Grok 4.6 scored 1,753 on GDPVal AA v2 compared with 1,728 for GPT 5.6 Sol Max
  6. Context github.com: Opus 5 and Antigravity Gemini 3.1 Pro (denoted by *) results updated August 2 2026
  7. Context github.com: Headline results on OfficeQA Pro (N=133), followed by OfficeQA Pro V2 (N=90)
  8. Supports eesel.ai: charging $2/$6 per million tokens against Sol's $5/$30
  9. Supports eesel.ai: Claude Opus 5 is $5/$25

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.