Subscribe

The trick: Borrowed Engine

Inherent says its 27-billion-parameter model beats GPT-5.5.

Its own paper says the model calls GPT-5.5 to do the work.

Issue 1625 August 20266 receipts4 min

Faraday, a 27-billion-parameter 'AI Scientist' from Inherent Labs, outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers.

Before you read on. Your call?

The "AI Scientist" that outperformed two frontier labs turns out to have one of them running inside it.

27BFaraday's advertised parameter count
73%share of in-distribution Replica tasks Faraday beats Opus 4.8 and GPT-5.5 on
60%win rate on out-of-distribution tasks
310total Replica benchmark tasks

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Borrowed Engine: it beats the thing it is secretly calling. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“"Which parts of the pipeline are your 27B model, and which parts are GPT-5.5 doing the heavy lifting?"”

Receipts

  1. Refutes arxiv.org: Faraday is provided with a frontier coding agent to use as a tool. A wrapper script runs the Codex CLI non-interactively.
  2. Refutes arxiv.org: GPT-5.5 in the final stage and for evaluation
  3. Refutes aiweekly.co: Faraday is built on a 27-billion-parameter Qwen base and calls OpenAI's GPT-5.5 Codex for coding subtasks.
  4. Context superpowerdaily.com: The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison.
  5. Supports techcrunch.com: What was most interesting to us about this was not so much the result of beating those frontier agents
  6. Supports inherentlabs.ai: a 27B-parameter "AI scientist" agent that outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.