The trick: Moved Ruler
Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks.
Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.
frontier AI cannot do professional investment analysis. On Samaya AI's FrontierFinance, a public benchmark of 220 expert queries, every frontier model scores below 50%, with Claude Fable 5 best at 49.2%, GPT-5.6 Sol at 46.8%, and Samaya's own system ahead of all of them at 56%.
Before you read on. Your call?
TRUE, BUT
49.2%
the ceiling is real and it is not a fluke. A separate benchmark from Vals AI, which builds evaluations rather than investment products but still profits when models fall short, reaches the same place: its top model, GPT-5.5, hits roughly 52%, with the frontier clustered in the high-40s to low-50s. Two separate teams, two methods, one wall. What the headline underplays is the shape of the sales floor around it. FrontierFinance is Samaya's own benchmark and Samaya's own system tops it, which is the oldest move in the book, a vendor grading an exam its product is built to pass. Even that winning system clears only 56%, and on the hardest categories, Screening and Discovery and Sector and Macro, the best of everything reaches 33% and 39%. The ceiling is the story. So is the fact that the people ringing the bell are selling the ladder.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →
Receipts
- Supports arxiv.org:
Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost
- Context arxiv.org:
Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%.
- Supports prnewswire.com:
Anthropic's Claude Fable 5 was the best performing frontier model, scoring 49.2%, followed by GPT-5.6 Sol at 46.8%
- Supports prnewswire.com:
Samaya's AI system outperformed every model evaluated.
- Context prnewswire.com:
Investment use cases are uniquely hard for AI because being almost correct is still a loss.
- Refutes kucoin.com:
GPT-5.5 scored approximately 52% accuracy on the Vals AI Finance Agent v2 benchmark, making it the top-performing model
- Context vals.ai:
are the top three performers, separated by less than a point
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.