Subscribe

The trick: Self-Marked

The new best open model beat GPT and Claude in every test.

Try finding one you can verify.

Issue 26 August 20263 receipts3 min

Meet Qwen3.8-Max, our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released

Before you read on. Your call?

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Ask any model vendor one question: 'Which of these benchmark numbers appears on a leaderboard you don't run?' Then count the seconds of silence.”

Receipts

  1. Supports latent.space: Vals Index: #2 among open models, 66.1 score (matched Claude Opus 4.7)
  2. Refutes yottalabs.ai: every stronger claim than that currently traces back to Alibaba's own press materials
  3. Context glbgpt.com: the table mixes official model reports, system cards, public leaderboards, and Qwen's in-house evaluations

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.