The trick: Borrowed Engine
Inherent says its 27-billion-parameter model beats GPT-5.5.
Its own paper says the model calls GPT-5.5 to do the work.
Faraday, a 27-billion-parameter 'AI Scientist' from Inherent Labs, outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers.
Before you read on. Your call?
TRUE, BUT
27b wrapper
The "AI Scientist" that outperformed two frontier labs turns out to have one of them running inside it.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Borrowed Engine: it beats the thing it is secretly calling. You'll see it again. Learn to spot it →
Receipts
- Refutes arxiv.org:
Faraday is provided with a frontier coding agent to use as a tool. A wrapper script runs the Codex CLI non-interactively.
- Refutes arxiv.org:
GPT-5.5 in the final stage and for evaluation
- Refutes aiweekly.co:
Faraday is built on a 27-billion-parameter Qwen base and calls OpenAI's GPT-5.5 Codex for coding subtasks.
- Context superpowerdaily.com:
The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison.
- Supports techcrunch.com:
What was most interesting to us about this was not so much the result of beating those frontier agents
- Supports inherentlabs.ai:
a 27B-parameter "AI scientist" agent that outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.