Subscribe

The trick: Self-Marked

Anthropic built a model stronger than its flagship.

You cannot use it, test it, or check the number.

Issue 1016 August 20265 receipts4 min

Anthropic's August 2026 Risk Report reveals Model 2, an internal system 'somewhat more capable' than Mythos 5 (62.8% vs 50.3% on CoBench), which the company does not plan to release; the same report raises Anthropic's catastrophic-misalignment rating from very low to low.

Before you read on. Your call?

the disclosure is real transparency and the capability claim is a closed loop. The score comes from an internal benchmark, run internally, on a model nobody outside Anthropic can touch: no weights, no API, no third-party run. Anthropic itself lowers the claim's confidence in ink, saying it 'has not run all of its typical predeployment assessments and therefore has somewhat less confidence in its beliefs about the model's capabilities'. Meanwhile the risk-label change the headlines attach to this scary stronger model has a different stated cause: increased overall uncertainty from recent cyber-evaluation incident disclosures involving shipped models, not Model 2. The report even notes its concrete task evaluations have 'saturated', meaning the measuring sticks maxed out. A stronger secret model and a raised risk label are both in the report. The causal arrow between them is not.

62.8%MODEL 2 ON COBENCH
50.3%THE SHIPPED FLAGSHIP ON THE SAME INTERNAL TEST
0OUTSIDE RUNS: NO WEIGHTS

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“The 62.8 is an internal score on an internal test of a model nobody outside can run, and Anthropic itself says its assessment is incomplete. The risk label moved for a different reason than the stronger-model headline implies.”

Receipts

  1. Supports unite.ai: We do not currently have plans to release this model externally
  2. Context unite.ai: to reflect increased overall uncertainty
  3. Supports finance.biggo.com: scoring 62.8% on CoBench versus Mythos 5's 50.3%
  4. Refutes techi.com: has not run all of its typical predeployment assessments and therefore has somewhat less confidence in its beliefs about the model's capabilities
  5. Context techi.com: even though its underlying argument likely still supports

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.