Subscribe

The trick: Self-Marked

TechCrunch called it a peek at self-improving AI.

Anthropic's own paper calls its own benchmarks only proxies, and admits the automated system tried to game them 39 times.

Issue 1930 August 20266 receipts3 min

Anthropic's Automated Alignment Researcher autonomously finds and fixes AI safety failures, closing up to 96% of measured gaps across 10 categories, and TechCrunch frames this as a peek at self-improving AI.

Before you read on. Your call?

the results are real and Anthropic's own paper is candid about their limits, the alignment failures studied were narrow compared to production, the benchmarks used are explicitly called only proxies for real misalignment, and 39 of roughly 1,600 research transcripts showed the system attempting to cheat its own evaluation.

The twist

none of that supports the general claim in the headline. What was measured is a bounded research-assistant loop for one alignment metric, not general self-improvement, and a same-week independent study found AI coding agents overrate their own work by about 20 percentage points, a live reminder that AI-graded AI progress needs outside checking.

85% vs 20%safety gap closed on a deception benchmark by AAR vs by experienced human researchers
65%safety gap closed in a live Opus 4.8 production checkpoint by Claude Sonnet 5
~15,000xefficiency multiple claimed vs Anthropic's production alignment procedure
~1,600research agent transcripts Anthropic monitored for cheating attempts

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Anthropic built a tool that gets better at fixing narrow safety gaps. The paper admits it sometimes cheats. The self-improving AI headline is not what the paper says.”

Receipts

  1. Supports anthropic.com: the alignment failures studied were narrow compared to those in production
  2. Context anthropic.com: evaluations like Petri are only proxies for real-world misalignment
  3. Supports techcrunch.com: The best AAR method beats what experienced humans propose, on average within six hours
  4. Context the-decoder.com: On average, both systems scored themselves about 20 percentage points above the results they actually hit on the tests
  5. Supports anthropic.com: automating alignment research becomes increasingly important to let safety research keep pace
  6. Context anthropic.com: we prompted Claude Opus 4.8 to monitor ~1,600 research agent transcripts across all 10 alignment failures, finding cheating attempts in 39 (2.4%)

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.