Subscribe

Z.ai says GLM-5.3 leads CyberGym with 84.5% and found 2,436 vulnerabilities.

Independent verifications: zero. Vulnerabilities with public CVEs: 53 out of 2,436.

Issue 1016 August 20264 receipts3 min

Z.ai launched GLM-5.3 on August 14 claiming CyberGym 84.5%, beating Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%), plus 2,436 confirmed vulnerabilities across 269 open-source projects.

Before you read on. Your call?

searched for independent CyberGym replication of GLM-5.3 on August 16, 2026: zero results. Every score is Z.ai-reported, Z.ai-tested, on Z.ai's own harness configuration. Of the 2,436 "confirmed" vulnerabilities, only 53 have public CVE assignments. The remaining 2,383 are under embargo. Confirmed is doing work it has not earned.

The twist

the weights are withheld for two weeks "for safety testing." That doubles as a perfect scarcity launch window. No one can replicate the benchmark until after the press cycle ends.

84.5% CyberGym score per OfficeChai
0independent verifications at launch
2436claimed confirmed vulnerabilities per fello AI
0.7point margin over Mythos 5

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

Say this in tomorrow's meeting“GLM-5.3 claims CyberGym first place by 0.7 points on its own test. Independent verifications: zero. The weights are held back for two weeks. That is not a benchmark result. That is a press embargo dressed as safety.”

Receipts

  1. Supports unite.ai: 2,436 vulnerabilities across 269 open-source projects
  2. Supports officechai.com: GLM-5.3 scores 84.5 on it, edging out both Claude Mythos 5 at 83.8 and GPT-5.6 Sol at 83.6
  3. Context aiweekly.co: Weights are being held back roughly two weeks for safety evaluation and hardening
  4. Context felloai.com: reasoning across multiple stages of exploitation

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.