The trick: Self-Marked
Alibaba's new small model tops a benchmark called QwenSWEBench.
Read the name again.
Alibaba's Qwen3.8-27B, open weights under Apache 2.0, posts frontier-tier agentic scores for a 27B model (61.7 SWE-bench Pro, 73.0 Terminal-Bench 2.1, 84.3 OSWorld-Verified) and beats Meta's Muse Glimmer 30B on the published comparison rows.
Before you read on. Your call?
TRUE, BUT
79.0
every number on the card was measured by Qwen. The widest margin, 79.0, lands on QwenSWEBench, a benchmark with the vendor's name in the title, and Muse Glimmer's results are missing entirely from several of the harder benchmarks Alibaba ran. Launch-day reviewers found no independent reproduction of any Qwen3.8-27B score. The one-gaming-GPU framing also needs fine print: the official BF16 build wants about 80GB with the KV cache at native context; the 24GB story is a third-party quant at moderate context. The weights are genuinely downloadable, so this one is checkable. It just has not been checked.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →
Receipts
- Supports officechai.com:
Qwen3.8-27B scores 73.0 against Muse Glimmer's 51.7
- Context officechai.com:
Muse Glimmer's results are missing entirely from several of the harder benchmarks Alibaba ran, including NL2Repo-Bench, DeepSWE 1.1, JobBench, and LiveCodeBench v6
- Refutes yottalabs.ai:
These are Alibaba's own model-card numbers, not independent replications, so the usual rule applies: validate before you build on them.
- Refutes kingy.ai:
Every launch score comes from Qwen. Several benchmarks are in-house, corrected or modified.
- Context kingy.ai:
A 24GB GPU is a plausible target for a four-bit quant at moderate context, not for BF16, FP8 or the full 262K window.
- Context kingy.ai:
Treat as an 80GB-class or multi-GPU deployment; full context still needs engine-specific proof
- Supports huggingface.co:
we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date
- Context huggingface.co:
QwenSWEBench: In-house coding benchmark for evaluating models' software engineering capabilities.
- Context officechai.com:
Qwen posts 79.5 to Muse Glimmer's 77.0
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.