Subscribe

The trick: Zero Underneath

Sol Ultrafast finished Humanity's Last Exam in 11 hours.

Whether it is still the same Sol remains unexamined.

Issue 814 August 20267 receipts3 min

OpenAI and Cerebras say GPT-5.6 Sol in Ultrafast mode hits up to 750 output tokens per second, 11x faster than Fable 5, and blitzed all 2,500 Humanity's Last Exam questions in 11 hours 11 minutes against Fable 5's 78 hours 27 minutes, with no quality compromise.

Before you read on. Your call?

the silicon is plausible and the framing is theater. Answering 2,500 independent questions is an embarrassingly parallel workload, so the wall-clock race measures cluster scale, not model speed, and per-question latency is not disclosed. Neither company states that Ultrafast performs identically to regular Sol, a sentence they would shout if they could write it, and the industry's record on 'no quality loss' serving modes is poor. No price, no context limit, no configuration. The fine print concedes results may vary by workload and configuration.

750TOK/S
62.2WHAT API SOL MEASURES ON THE OPEN API
0PRICES

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Zero Underneath: the headline number has nothing behind it. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“The fast is probably real, the parity is asserted. Until there is per-question latency, a price, and a full eval suite on Ultrafast, treat it as a different product.”

Receipts

  1. Supports cerebras.ai: up to 750 output tokens per second
  2. Supports cerebras.ai: 11x faster than Fable 5
  3. Supports cerebras.ai: without any quality compromise
  4. Refutes cerebras.ai: may vary depending on workload, configuration
  5. Refutes news.ycombinator.com: answering 2,500 independent questions is an embarrassingly parallel workload, all it needs is scale out.
  6. Refutes news.ycombinator.com: I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops.
  7. Context web.archive.org: 62.2 Output tokens per second

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.