The trick: Narrowed Superlative
Claude Fable 5 really is number one on the hardest AI leaderboards.
It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.
Claude Fable 5, Anthropic's most capable public model, is the new benchmark leader, topping the Artificial Analysis Intelligence Index and Humanity's Last Exam and setting the highest score to date on AA-Omniscience, the knowledge and hallucination benchmark.
Before you read on. Your call?
TRUE, BUT
55.5%
the ranking is real. Fable 5 sits at number one on these boards on independent evaluation, which is not nothing. What the headline hides is what the numbers mean. AA-Omniscience runs from minus 100 to 100, where zero means as many right as wrong and most frontier models score below zero, so a leading 43 is a genuine jump and still a long way from anything a layperson would call omniscient. Humanity's Last Exam was built so models fail it, seeding only questions that already stumped the best AIs, so a rank of number one at 55.5% is a lead, not a grade. And the leaderboard entry is a specific named configuration: the record holder is the Adaptive Reasoning, Max Effort, Opus 4.8 Fallback setup, and Anthropic has not disclosed how the number would move without the fallback. Number one is a ranking, not a report card.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Narrowed Superlative: first or fastest, inside a quietly narrowed category. You'll see it again. Learn to spot it →
Receipts
- Supports artificialanalysis.ai:
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) scores the highest on Humanity's Last Exam with a score of 55.5%
- Context artificialanalysis.ai:
Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
- Supports artificialanalysis.ai:
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) scores the highest on AA-Omniscience with a score of 43
- Supports siliconreport.com:
Claude Fable 5 is the most capable model Anthropic sells to the general public
- Context siliconreport.com:
Anthropic says more than 95% of Fable 5 sessions involve no fallback whatsoever
- Refutes the-decoder.com:
There's a big gulf between what it means to take an exam and what it means to be a practicing physicist and researcher.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.