Subscribe

The trick: Lab Not Field

Anthropic studied whether you read permission prompts.

You don't. So today it stopped showing them.

Issue 915 August 20264 receipts4 min

auto mode matched or outperformed manual review on every measure Anthropic tested, catching 89% of planted dangerous commands against 13.6% for humans, so from August 14 it is the default for Pro, Max and Team users.

Before you read on. Your call?

the numbers come from Anthropic's own unreviewed study, the winning comparison is against fatigued humans in a lab scenario, and Anthropic's engineering post calls a different figure 'the honest number': a 17% false-negative rate on real dangerous actions, from a sample of just 52. Anthropic itself still recommends human review for high-risk changes.

17%REAL OVEREAGER ACTIONS THE SHIPPED PIPELINE LETS THROUGH
52SAMPLE SIZE BEHIND THAT HONEST NUMBER
13.6%HUMAN CATCH RATE IN THE PLANTED-COMMAND TEST

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“'What is the false-negative rate on real incidents rather than planted ones?' Anthropic printed it: 17%, from 52 examples. Ask why the press release quotes the other number.”

Receipts

  1. Supports theregister.com: In the controlled study, testers caught a deliberately inserted dangerous command just 13.6 percent of the time. Auto mode blocked 89 percent of the same commands.
  2. Refutes anthropic.com: The 17% false-negative rate on real overeager actions is the honest number.
  3. Context helpnetsecurity.com: Auto mode reduces risk for most users but does not eliminate it because it relies on AI to judge whether actions are safe.
  4. Context arxiv.org: The end-to-end false negative rate is 81.0% (95% CI: 73.8%-87.4%), substantially higher than the 17% reported on production traffic, reflecting a fundamentally different workload rather than a contradiction

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.