Subscribe

The trick: Cherry-Picked Slice

Nvidia's agent just went 100% on a benchmark built to resist that.

The brain doing the reasoning is Anthropic's, and it scores 30% alone.

Issue 1625 August 20265 receipts4 min

NVIDIA AVO achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels, demonstrating a frontier-level general-purpose architecture for long-horizon autonomous agents.

Before you read on. Your call?

A perfect score on the public half of a test that has never been beaten on the half that counts.

100.00RHAE score
0systems
30.2%Claude Opus 5 bare baseline score at high reasoning effort - the model powering AVO's reasoning
6,624environment actions AVO used

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“"Show me the private-set score, or tell me why there isn't one yet."”

Receipts

  1. Refutes developer.nvidia.com: They are not results on the semi-private or fully private competition sets.
  2. Refutes developer.nvidia.com: This should not be interpreted as a controlled ablation: the two systems differ in agent backend, observation representation, memory, context management
  3. Refutes eneralabs.com: The private-set benchmark, which withholds environments not available in the public set, remains unsolved.
  4. Refutes cryptobriefing.com: Without private set results, the 100% score is impressive but incomplete as a measure of general reasoning ability.
  5. Context cryptobriefing.com: AVO exists as a research demonstration for now, not a product you can buy or integrate.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.