The trick: Narrowed Superlative
Microsoft's new model goes toe-to-toe with the Claude that was champion in June.
It is August.
Microsoft's first in-house reasoning model matches Claude Opus 4.6 on SWE-Bench Pro, reaches 97.0% on AIME 2025, and beat Claude Sonnet 4.6 in blind human preference evaluations, all while running 35B active parameters at a mid-weight price.
Before you read on. Your call?
TRUE, BUT
97.0%
every comparison targets the previous Claude generation. Opus 4.6 was the frontier in spring; Opus 5 shipped July 24 and Fable 5 sits above it, and the model Microsoft beat on preference, Sonnet 4.6, is the mid-tier of that older line. The scores are self-reported from a vendor preprint, the human eval was commissioned by Microsoft from its rating partner Surge, an independent aggregator has not confirmed the flagship AIME figure, and Artificial Analysis lists no page for the model at all. The claims were minted at Build in June; the Foundry rollout re-airs them unchanged, two Claude generations later.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Narrowed Superlative: first or fastest, inside a quietly narrowed category. You'll see it again. Learn to spot it →
Receipts
- Supports microsoft.ai:
Despite this, our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro.
- Supports microsoft.ai:
MAI-Thinking-1 reaches 97.0% on AIME 2025, and 94.5% on AIME 2026, showing strong mathematical and scientific reasoning for its weight class.
- Context microsoft.ai:
The evaluation spanned 1,276 tasks across a wide variety of use cases in both single-turn and multi-turn conversations
- Context microsoft.ai:
users preferred MAI-Thinking-1 over Claude Sonnet 4.6
- Refutes techjacksolutions.com:
Benchmark scores are self-reported and one independent aggregator hasn't confirmed the flagship AIME 2025 figure.
- Context techtimes.com:
First In-House Reasoning Model, Trained Without OpenAI Data
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.