Subscribe

The trick: Cherry-Picked Slice

Meta says its new model beat two rivals across half the benchmarks.

Half.

Issue 511 August 20264 receipts3 min

Meta released Muse Glimmer, a 30B open-weight model, and said it outperformed the similarly sized Gemma4-31B and Qwen3.6-27B across half the benchmarks.

Before you read on. Your call?

'half the benchmarks' means it also lost half. That is the number Meta chose to lead with. Independent testing by Artificial Analysis puts Qwen3.6-27B ahead overall, 38 to 35 on its Intelligence Index, and clocks Glimmer hallucinating at 82% versus Qwen's 49%.

The twist

this is not an open version of Meta's most powerful model. That is Muse Spark 1.2, whose weights are promised 'soon', with no date. Glimmer is the real news anyway: after more than a year of closed releases, Meta shipped actual Apache 2.0 weights that run on a single consumer GPU. Good story. Bad spin on top of it.

HALFBENCHMARKS WON
82%GLIMMER HALLUCINATION RATE
49%QWEN3.6-27B SAME EVAL
35GLIMMER

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“'Beat them on which benchmarks, and who ran the test?' Glimmer 'won half' on Meta's own board and lost the class on someone else's.”

Receipts

  1. Supports siliconangle.com: outperformed the comparably-sized Gemma4-31B and Qwen3.6-27B across half the benchmarks
  2. Supports research.meta.ai: Muse Glimmer performs strongly for its size class on several widely used LLM benchmarks
  3. Refutes theregister.com: As usual, take these claims with a grain of salt
  4. Refutes artificialanalysis.ai: an 82% hallucination rate (Qwen3.6 27B: 49%)

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.