The trick: Cherry-Picked Slice
A stealth model beat Claude and GPT on a coding benchmark.
The sample size was ten questions.
A tester's viral post said the free stealth model Ox Alpha scored 80% on the DeepSWE coding benchmark, beating GPT-5.6 Sol's 52% and Claude Fable 5's 65%, and Coin Bureau broadcast it to a much larger audience as a mystery model beating the leaders.
Before you read on. Your call?
TRUE, BUT
80% to 58%
the tester had run 10 of the benchmark's 113 tasks. He ran the rest the next day and landed at 63%. An independent lab ran all 113 and got 58.4%, calling the 80% figure completely incorrect.
The twist
the model's maker, Zhipu AI, has since confirmed its identity as GLM-5.3-Flash, and on a more rigorous multi-trial benchmark the base GLM-5.3 model (not confirmed as the same Flash variant) loses to GPT-5.6 Sol on a single attempt but ties or leads it once retries are allowed.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →
Receipts
- Supports x.com:
gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80%
- Refutes x.com:
Actual DeepSWE run on the ox alpha mystery model is done. Ended at ~63% NOT the 80% my first subset test got
- Refutes stealthmodelwatch.online:
The rumored ~80% pass rate is completely incorrect
- Context together.ai:
Sol edges ahead: 72.7% pass@1 to GLM-5.3's 69.0%
- Supports x.com:
BREAKING: A mysterious new AI model called "Ox Alpha" is reportedly BEATING Claude Fable 5 and GPT-5.6 Sol at coding, and nobody knows who built it
- Context vktr.com:
On August 26, Z.ai confirmed authorship of the model, vindicating the fingerprinting. The model's official name is GLM-5.3-Flash
- Context together.ai:
It ties Sol at pass@2 (81.1 vs. 81.0) and leads pass@4 (87.6% vs. 85.8%)
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.