Subscribe

The trick: Cherry-Picked Slice

The '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.

The Nature paper is real and its warning is serious. The number people quote from it is an average across weak targets, and the appendix that debunks the headline is in the same paper.

Issue 814 August 20263 receipts3 min

Reasoning models now jailbreak frontier LLMs autonomously at a 97% success rate

Before you read on. Your call?

the 97.14% figure comes from Hagendorff, Derner and Oliver's Nature Communications study, and it is an average across all attacker-target combinations, including old and weak targets. The paper's own data says the most resistant target, Claude 4 Sonnet, took the top harm score on just 2.86% of items, and attacker success ranged from 12.86% to 90%. Real alignment finding, real warning. The single scary number flattens all of it.

2.86%HARM RATE AGAINST THE MOST RESISTANT TARGET
97.14%THE AGGREGATE
12.86%SUCCESS RATE OF THE WEAKEST ATTACKER

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“'97% against which target?' The paper names them. Against Claude 4 Sonnet the harm rate was 2.86%. Against old DeepSeek-V3 it was 90%. The average is not the story.”

Receipts

  1. Supports sqmagazine.co.uk: Multi-turn jailbreaks hit 97% success on frontier LLMs.
  2. Refutes nature.com: Claude 4 Sonnet is by far the most resistant model, receiving the highest harm score in only a fraction of benchmark items across adversarial models (2.86%)
  3. Context web.archive.org: the persuasive capabilities of large reasoning models (LRMs) simplify and scale jailbreaking, converting it into an inexpensive activity accessible to non-experts

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.