The trick: Announced Not Shipped
Grok 4.6 posts a 1753 and a number one.
One is a real benchmark. The other is a tweet.
Musk says Grok 4.6 reaches 1753 ELO and is number one on Databricks' OfficeQA benchmark, at half the price of rivals.
Before you read on. Your call?
TRUE, BUT
1753
the 1753 is real, and it comes from Artificial Analysis, not just Musk. But it is one benchmark, GDPval-AA v2, where Grok sits behind Claude Opus 5 and inside overlapping confidence intervals with Fable 5 and Qwen. On the overall Intelligence Index it scores 61, tied with GPT-5.6 Sol for third, behind Opus at 63 and Fable at 62.
The twist
the number one Musk tweeted is an OfficeQA Pro V2 run with Databricks' Genie harness, and in one analysis's words that number 'is not in SpaceXAI's published table.' xAI's model card does show a number one on the older OfficeQA Pro v1, 63.2 percent, but xAI ran that itself, and Databricks' own leaderboard does not list Grok 4.6 at all. The part that actually holds is the boring part: Grok 4.6 is cheap, about $0.84 a task, with list prices of $2 in and $6 out per million tokens against $5 and $25 for Opus 5.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Announced Not Shipped: claimed, never released. You'll see it again. Learn to spot it →
Receipts
- Context artificialanalysis.ai:
achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5
- Context artificialanalysis.ai:
It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62)
- Refutes explainx.ai:
That number is not in SpaceXAI's published table.
- Context explainx.ai:
Treat it as a founder claim until Databricks or a third party reproduces it.
- Supports cryptobriefing.com:
Grok 4.6 scored 1,753 on GDPVal AA v2 compared with 1,728 for GPT 5.6 Sol Max
- Context github.com:
Opus 5 and Antigravity Gemini 3.1 Pro (denoted by *) results updated August 2 2026
- Context github.com:
Headline results on OfficeQA Pro (N=133), followed by OfficeQA Pro V2 (N=90)
- Supports eesel.ai:
charging $2/$6 per million tokens against Sol's $5/$30
- Supports eesel.ai:
Claude Opus 5 is $5/$25
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.