The trick: Self-Marked
OpenAI's president says we're in the AGI era.
The benchmark's own inventor scored the same model 37 points lower.
OpenAI president Greg Brockman said it's not unreasonable to feel that we are now in the AGI era, pointing to Astra's 99.9% on ARC-AGI-3.
Before you read on. Your call?
TRUE, BUT
37-point drop
ARC Prize, the outfit that built the benchmark, ran the same model on its neutral harness and got 62.7%, a 37-point drop, and said outright it is not claiming AGI.
The twist
on the Artificial Analysis Intelligence Index, a broader aggregate less prone to a single-harness trick, Astra ties its own six-month-old predecessor and loses to Claude Fable 5.1, while charging 2.5x more per token to do it.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →
Receipts
- Supports fortune.com:
It's not unreasonable to feel that we are now in the AGI era
- Refutes arcprize.org:
GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness
- Refutes arcprize.org:
we are not claiming that it is AGI
- Context requesty.ai:
Astra is good, but maybe we should calm down with the hype
- Context emergent.sh:
On the Artificial Analysis Intelligence Index, which aggregates reasoning, knowledge, and coding evaluations into one score, Astra landed at 61.2. That ties GPT-5.6 Sol at 60.9 and trails Claude Fable 5.1 at 65.7.
- Context emergent.sh:
On Humanity's Last Exam with tools, Astra's 57.2% clearly trails Fable 5.1's 65.0%.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.