The trick: Self-Marked
Anthropic built a model stronger than its flagship.
You cannot use it, test it, or check the number.
Anthropic's August 2026 Risk Report reveals Model 2, an internal system 'somewhat more capable' than Mythos 5 (62.8% vs 50.3% on CoBench), which the company does not plan to release; the same report raises Anthropic's catastrophic-misalignment rating from very low to low.
Before you read on. Your call?
TRUE, BUT
62.8%
the disclosure is real transparency and the capability claim is a closed loop. The score comes from an internal benchmark, run internally, on a model nobody outside Anthropic can touch: no weights, no API, no third-party run. Anthropic itself lowers the claim's confidence in ink, saying it 'has not run all of its typical predeployment assessments and therefore has somewhat less confidence in its beliefs about the model's capabilities'. Meanwhile the risk-label change the headlines attach to this scary stronger model has a different stated cause: increased overall uncertainty from recent cyber-evaluation incident disclosures involving shipped models, not Model 2. The report even notes its concrete task evaluations have 'saturated', meaning the measuring sticks maxed out. A stronger secret model and a raised risk label are both in the report. The causal arrow between them is not.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →
Receipts
- Supports unite.ai:
We do not currently have plans to release this model externally
- Context unite.ai:
to reflect increased overall uncertainty
- Supports finance.biggo.com:
scoring 62.8% on CoBench versus Mythos 5's 50.3%
- Refutes techi.com:
has not run all of its typical predeployment assessments and therefore has somewhat less confidence in its beliefs about the model's capabilities
- Context techi.com:
even though its underlying argument likely still supports
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.