Subscribe

The trick: Moved Ruler

Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks.

Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.

Issue 1319 August 20267 receipts4 min

frontier AI cannot do professional investment analysis. On Samaya AI's FrontierFinance, a public benchmark of 220 expert queries, every frontier model scores below 50%, with Claude Fable 5 best at 49.2%, GPT-5.6 Sol at 46.8%, and Samaya's own system ahead of all of them at 56%.

Before you read on. Your call?

the ceiling is real and it is not a fluke. A separate benchmark from Vals AI, which builds evaluations rather than investment products but still profits when models fall short, reaches the same place: its top model, GPT-5.5, hits roughly 52%, with the frontier clustered in the high-40s to low-50s. Two separate teams, two methods, one wall. What the headline underplays is the shape of the sales floor around it. FrontierFinance is Samaya's own benchmark and Samaya's own system tops it, which is the oldest move in the book, a vendor grading an exam its product is built to pass. Even that winning system clears only 56%, and on the hardest categories, Screening and Discovery and Sector and Macro, the best of everything reaches 33% and 39%. The ceiling is the story. So is the fact that the people ringing the bell are selling the ladder.

49.2%BEST FRONTIER MODEL
56%SAMAYA'S OWN SYSTEM ON SAMAYA'S OWN BENCHMARK
33% / 39%WHERE EVEN THE BEST SYSTEM LANDS ON THE HARDEST USE CASES
~52%TOP MODEL ON THE SEPARATE VALS AI FINANCE BENCHMARK

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Two independent finance benchmarks put the best AI models under the low-50s on real analyst tasks, so the capability ceiling is real. But FrontierFinance is Samaya's own benchmark, Samaya's own system wins it at 56%, and even that winner clears only just over half. The wall is genuine. The scoreboard is marketing.”

Receipts

  1. Supports arxiv.org: Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost
  2. Context arxiv.org: Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%.
  3. Supports prnewswire.com: Anthropic's Claude Fable 5 was the best performing frontier model, scoring 49.2%, followed by GPT-5.6 Sol at 46.8%
  4. Supports prnewswire.com: Samaya's AI system outperformed every model evaluated.
  5. Context prnewswire.com: Investment use cases are uniquely hard for AI because being almost correct is still a loss.
  6. Refutes kucoin.com: GPT-5.5 scored approximately 52% accuracy on the Vals AI Finance Agent v2 benchmark, making it the top-performing model
  7. Context vals.ai: are the top three performers, separated by less than a point

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.