Subscribe

The trick: Self-Marked

DeepSeek's chart said its flagship jumped 49.9 points.

The referee showed up and moved the index by one.

Issue 915 August 20267 receipts3 min

DeepSeek's official flagship endpoint now serves V4-Pro-0813, and the company's chart shows agentic scores exploding: DeepSWE from 12.8 to 62.7, CyberGym from 52.7 to 83.3, Terminal Bench 2.1 at 87.9.

Before you read on. Your call?

at launch, every number was DeepSeek grading DeepSeek against its own retired preview. Within a day the independent record filled in, and it reads smaller: Artificial Analysis scores the model 53 on its Intelligence Index, one point above DeepSeek's own cheaper Flash, and measures Terminal-Bench v2.1 at 79, well under the chart's 87.9 and ten points behind Claude Opus 5. The direction is real, the drama is not. SCMP adds that developers are underwhelmed and disappointed in the pricing, even as cybersecurity researchers are impressed.

79TERMINAL-BENCH 2.1
87.9SAME BENCHMARK
+1INTELLIGENCE INDEX MOVE OVER DEEPSEEK'S OWN FLASH

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“The vendor chart said 87.9. The independent run says 79. The composite index moved one point.”

Receipts

  1. Supports globaltimes.cn: Chinese artificial intelligence (AI) company DeepSeek officially unveiled the official version of DeepSeek V4 Pro on Thursday, featuring significantly enhanced agent capabilities
  2. Context dsv4pro.novcog.us.com: Terminal Bench 2.1 72.1 87.9 +15.8 CyberGym 52.7 83.3 +30.6 DeepSWE 12.8 62.7 +49.9
  3. Context dsv4pro.novcog.us.com: As of 12 August 2026, there is not one third-party benchmark result for DeepSeek-V4-Pro-0813 that this site could find. Every number on this page is vendor-reported or analyst-reported; none is independently reproduced.
  4. Refutes officechai.com: On Terminal-Bench v2.1, which tests agentic terminal workflows, DeepSeek V4 Pro scores 79%, a full ten points behind Claude Opus 5's 89%.
  5. Refutes officechai.com: Model scores one point higher than Flash on Artificial Analysis Intelligence Index
  6. Context artificialanalysis.ai: 53 Artificial Analysis Intelligence Index
  7. Context scmp.com: leaving some developers underwhelmed by its overall capabilities and disappointed in its pricing

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.