The trick: Cherry-Picked Slice
The '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.
The Nature paper is real and its warning is serious. The number people quote from it is an average across weak targets, and the appendix that debunks the headline is in the same paper.
Reasoning models now jailbreak frontier LLMs autonomously at a 97% success rate
Before you read on. Your call?
TRUE, BUT
2.86%
the 97.14% figure comes from Hagendorff, Derner and Oliver's Nature Communications study, and it is an average across all attacker-target combinations, including old and weak targets. The paper's own data says the most resistant target, Claude 4 Sonnet, took the top harm score on just 2.86% of items, and attacker success ranged from 12.86% to 90%. Real alignment finding, real warning. The single scary number flattens all of it.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →
Receipts
- Supports sqmagazine.co.uk:
Multi-turn jailbreaks hit 97% success on frontier LLMs.
- Refutes nature.com:
Claude 4 Sonnet is by far the most resistant model, receiving the highest harm score in only a fraction of benchmark items across adversarial models (2.86%)
- Context web.archive.org:
the persuasive capabilities of large reasoning models (LRMs) simplify and scale jailbreaking, converting it into an inexpensive activity accessible to non-experts
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.