Subscribe

The trick: Cherry-Picked Slice

A stealth model beat Claude and GPT on a coding benchmark.

The sample size was ten questions.

Issue 1728 August 20267 receipts3 min

A tester's viral post said the free stealth model Ox Alpha scored 80% on the DeepSWE coding benchmark, beating GPT-5.6 Sol's 52% and Claude Fable 5's 65%, and Coin Bureau broadcast it to a much larger audience as a mystery model beating the leaders.

Before you read on. Your call?

the tester had run 10 of the benchmark's 113 tasks. He ran the rest the next day and landed at 63%. An independent lab ran all 113 and got 58.4%, calling the 80% figure completely incorrect.

The twist

the model's maker, Zhipu AI, has since confirmed its identity as GLM-5.3-Flash, and on a more rigorous multi-trial benchmark the base GLM-5.3 model (not confirmed as the same Flash variant) loses to GPT-5.6 Sol on a single attempt but ties or leads it once retries are allowed.

80%Ben Davis's original viral small-sample score for Ox Alpha on DeepSWE
63%Ben Davis's own full 113-task re-run score
58.4%independent StealthModelWatch full 113-task run
72.7%Together AI's rigorous 4-trial pass@1 for GPT-5.6 Sol on DeepSWE

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Ten questions is a party trick, not a benchmark.”

Receipts

  1. Supports x.com: gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80%
  2. Refutes x.com: Actual DeepSWE run on the ox alpha mystery model is done. Ended at ~63% NOT the 80% my first subset test got
  3. Refutes stealthmodelwatch.online: The rumored ~80% pass rate is completely incorrect
  4. Context together.ai: Sol edges ahead: 72.7% pass@1 to GLM-5.3's 69.0%
  5. Supports x.com: BREAKING: A mysterious new AI model called "Ox Alpha" is reportedly BEATING Claude Fable 5 and GPT-5.6 Sol at coding, and nobody knows who built it
  6. Context vktr.com: On August 26, Z.ai confirmed authorship of the model, vindicating the fingerprinting. The model's official name is GLM-5.3-Flash
  7. Context together.ai: It ties Sol at pass@2 (81.1 vs. 81.0) and leads pass@4 (87.6% vs. 85.8%)

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.