Subscribe

The trick: Lab Not Field

The pitch is that frontier models can generate genuine research ideas.

A new blind benchmark handed seven of them a paper's reference list, scrubbed of anything they could have memorized, and asked for the paper's core idea. They got it 3 to 15 percent of the time.

Issue 1420 August 20266 receipts4 min

AI is becoming a scientific research partner, able to reason over the literature and generate genuine new ideas. Labs sell this constantly, from AI accelerating research to models proposing hypotheses.

Before you read on. Your call?

a benchmark called Reconstruction, posted to arXiv on August 17, tests the cleanest version of that claim. It gives a model only a paper's pre-publication bibliography, with the seed paper and all contemporaneous or future literature withheld and reference IDs anonymized, then asks it to propose the paper's actual core idea, which an independent model judge scores against the held-out truth. The anti-leakage design means a high score cannot come from having memorized the answer. Across 643 papers in six scientific domains, seven frontier models landed at approximately 3 to 15 percent. A reference-only multi-agent pipeline, cross-model review plus a Swiss tournament over competing hypotheses, raised that to roughly 23 to 42 percent, a genuine 2.4x lift that still leaves the majority of ideas unrecovered. Reproducing the paper's specific hypothesis from a reading list is the part the models mostly cannot, with the caveat that a model proposing a different but valid idea would also score zero here, so this measures recovery of a known answer rather than open idea generation.

3-15%SINGLE FRONTIER MODEL MATCH RATE ON RECONSTRUCTION
23-42%BEST MULTI-AGENT PIPELINE
2.4xLIFT FROM CROSS-MODEL REVIEW PLUS TOURNAMENT
643PAPERS ACROSS 6 DOMAINS

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“A contamination-blind benchmark gave seven frontier models a paper's reference list and asked for its core idea. Single models scored 3 to 15 percent, and the best multi-agent pipeline reached 23 to 42 percent, a real 2.4x lift that still misses most. AI is strong at synthesizing the literature and weak at recovering the specific hypothesis a given paper landed on.”

Receipts

  1. Refutes arxiv.org: Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%).
  2. Supports arxiv.org: Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline.
  3. Context arxiv.org: We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature
  4. Refutes techtimes.com: Solo scores of three to fifteen percent on a contamination-resistant task
  5. Context techtimes.com: tests seven frontier models against 643 papers across six scientific domains
  6. Context emergentmind.com: combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.