Subscribe

The trick: Lab Not Field

NVIDIA built a model that talks four times faster.

The work arrives 30 percent sooner.

Issue 713 August 20264 receipts3 min

Nemotron 3.5 Lightning, NVIDIA's new 30B agent model with 3B active parameters, generates output up to 4x faster than similar-sized models and wins the accuracy-speed Pareto frontier.

Before you read on. Your call?

three sentences after the 4x, NVIDIA's own blog concedes that agent efficiency comes down to completed work, not token speed, and its own PinchBench chart shows tasks finishing 30% faster than Qwen3.6 35B. The 4x is 'up to', measured by the vendor, and Artificial Analysis's page for the model lists no output speed at all. A 4x engine revs; a 30% car arrives. Both numbers are NVIDIA grading NVIDIA.

4xOUTPUT SPEED
30%FASTER THE ACTUAL TASKS FINISH
86%PINCHBENCH ACCURACY

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Tokens per second is the engine revving. Task completion is the car arriving. NVIDIA's own chart says 30 percent.”

Receipts

  1. Supports developer.nvidia.com: Nemotron 3.5 Lightning combines strong intelligence with up to 4x output speed of similar-sized models, placing it on the accuracy-speed Pareto frontier for high-volume agent workloads.
  2. Context developer.nvidia.com: Agent efficiency ultimately comes down to how quickly a model completes useful work and not simply how fast it generates tokens. On PinchBench, Nemotron 3.5 Lightning reaches 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at similar accuracy.
  3. Refutes xenospectrum.com: the "up to 4x" figure is NVIDIA's own claim, and Artificial Analysis's current model page does not list output speed. When deploying, completion times need to be re-measured using one's own prompt lengths and concurrency levels.
  4. Context datacamp.com: NVIDIA claims it wins the accuracy-versus-speed Pareto frontier on the Artificial Analysis Intelligence Index, and reports up to 4x output speed of similar-sized models.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.