Subscribe

The trick: Rented Halo

Yes, researchers pulled the hidden reasoning out of Claude, GPT, and Gemini.

No, it was not the token counts, and it is already fixed.

Issue 612 August 20266 receipts3 min

a team found a vulnerability in every frontier lab's API that leaks models' hidden reasoning, and you could verify it by matching the billable thinking-token counts.

Before you read on. Your call?

the paper is real and demonstrated across Anthropic, OpenAI, and Google. But the mechanism in the viral version is wrong. The token count did not leak the reasoning. The attack replays a strong model's encrypted reasoning block into a weaker, less-guarded sibling model, which transcribes it in plain text. Token-count matching only confirmed the recovered trace was exact.

The twist

the biggest fact got left out of the thread. After the researchers disclosed it, all three providers deployed server-side mitigations. It is a real and clever result about portable reasoning blocks, reported as a live catastrophe after the fixes had already shipped.

3FRONTIER LABS DEMONSTRATED
315,320REASONING BLOCKS DECODED
0LABS THAT LEFT IT UNFIXED AFTER DISCLOSURE

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Rented Halo: the achievement is real, the drama around it is borrowed. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Real attack, wrong mechanism, already patched. It replayed reasoning into a weaker model. The token count only proved the copy was exact.”

Receipts

  1. Supports arxiv.org: we demonstrate across Anthropic, OpenAI, and Google
  2. Context arxiv.org: By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider
  3. Context arxiv.org: By decoding 315,320 reasoning blocks scraped from public repositories
  4. Refutes cybersecuritynews.com: deployed server-side mitigations
  5. Refutes simonwillison.net: All model providers acknowledged the receipt of our report
  6. Context news.ycombinator.com: nothing scientific here, they basically just figured out some real issues caused by bad engineering practice.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.