Subscribe

The trick: Self-Marked

Alibaba's new small model tops a benchmark called QwenSWEBench.

Read the name again.

Issue 915 August 20269 receipts4 min

Alibaba's Qwen3.8-27B, open weights under Apache 2.0, posts frontier-tier agentic scores for a 27B model (61.7 SWE-bench Pro, 73.0 Terminal-Bench 2.1, 84.3 OSWorld-Verified) and beats Meta's Muse Glimmer 30B on the published comparison rows.

Before you read on. Your call?

every number on the card was measured by Qwen. The widest margin, 79.0, lands on QwenSWEBench, a benchmark with the vendor's name in the title, and Muse Glimmer's results are missing entirely from several of the harder benchmarks Alibaba ran. Launch-day reviewers found no independent reproduction of any Qwen3.8-27B score. The one-gaming-GPU framing also needs fine print: the official BF16 build wants about 80GB with the KV cache at native context; the 24GB story is a third-party quant at moderate context. The weights are genuinely downloadable, so this one is checkable. It just has not been checked.

79.0QWENSWEBENCH
0INDEPENDENT REPRODUCTIONS AT LAUNCH
80GBWHAT THE OFFICIAL BF16 BUILD ACTUALLY WANTS

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“The weights are open and that is real. The scores are Qwen grading Qwen, the widest win is on Qwen's own benchmark, and nobody has reproduced a single number yet.”

Receipts

  1. Supports officechai.com: Qwen3.8-27B scores 73.0 against Muse Glimmer's 51.7
  2. Context officechai.com: Muse Glimmer's results are missing entirely from several of the harder benchmarks Alibaba ran, including NL2Repo-Bench, DeepSWE 1.1, JobBench, and LiveCodeBench v6
  3. Refutes yottalabs.ai: These are Alibaba's own model-card numbers, not independent replications, so the usual rule applies: validate before you build on them.
  4. Refutes kingy.ai: Every launch score comes from Qwen. Several benchmarks are in-house, corrected or modified.
  5. Context kingy.ai: A 24GB GPU is a plausible target for a four-bit quant at moderate context, not for BF16, FP8 or the full 262K window.
  6. Context kingy.ai: Treat as an 80GB-class or multi-GPU deployment; full context still needs engine-specific proof
  7. Supports huggingface.co: we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date
  8. Context huggingface.co: QwenSWEBench: In-house coding benchmark for evaluating models' software engineering capabilities.
  9. Context officechai.com: Qwen posts 79.5 to Muse Glimmer's 77.0

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.