The trick: Cherry-Picked Slice
Nvidia's agent just went 100% on a benchmark built to resist that.
The brain doing the reasoning is Anthropic's, and it scores 30% alone.
NVIDIA AVO achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels, demonstrating a frontier-level general-purpose architecture for long-horizon autonomous agents.
Before you read on. Your call?
TRUE, BUT
100.00 public-only
A perfect score on the public half of a test that has never been beaten on the half that counts.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →
Receipts
- Refutes developer.nvidia.com:
They are not results on the semi-private or fully private competition sets.
- Refutes developer.nvidia.com:
This should not be interpreted as a controlled ablation: the two systems differ in agent backend, observation representation, memory, context management
- Refutes eneralabs.com:
The private-set benchmark, which withholds environments not available in the public set, remains unsolved.
- Refutes cryptobriefing.com:
Without private set results, the 100% score is impressive but incomplete as a measure of general reasoning ability.
- Context cryptobriefing.com:
AVO exists as a research demonstration for now, not a product you can buy or integrate.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.