The trick: Lab Not Field
The pitch is that frontier models can generate genuine research ideas.
A new blind benchmark handed seven of them a paper's reference list, scrubbed of anything they could have memorized, and asked for the paper's core idea. They got it 3 to 15 percent of the time.
AI is becoming a scientific research partner, able to reason over the literature and generate genuine new ideas. Labs sell this constantly, from AI accelerating research to models proposing hypotheses.
Before you read on. Your call?
TRUE, BUT
3-15%
a benchmark called Reconstruction, posted to arXiv on August 17, tests the cleanest version of that claim. It gives a model only a paper's pre-publication bibliography, with the seed paper and all contemporaneous or future literature withheld and reference IDs anonymized, then asks it to propose the paper's actual core idea, which an independent model judge scores against the held-out truth. The anti-leakage design means a high score cannot come from having memorized the answer. Across 643 papers in six scientific domains, seven frontier models landed at approximately 3 to 15 percent. A reference-only multi-agent pipeline, cross-model review plus a Swiss tournament over competing hypotheses, raised that to roughly 23 to 42 percent, a genuine 2.4x lift that still leaves the majority of ideas unrecovered. Reproducing the paper's specific hypothesis from a reading list is the part the models mostly cannot, with the caveat that a model proposing a different but valid idea would also score zero here, so this measures recovery of a known answer rather than open idea generation.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →
Receipts
- Refutes arxiv.org:
Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%).
- Supports arxiv.org:
Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline.
- Context arxiv.org:
We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature
- Refutes techtimes.com:
Solo scores of three to fifteen percent on a contamination-resistant task
- Context techtimes.com:
tests seven frontier models against 643 papers across six scientific domains
- Context emergentmind.com:
combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.