The trick: Cherry-Picked Slice
GPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it.
Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.
GPT-5.6 Sol tops the ARC-AGI-2 leaderboard at 92.5%, clearing the benchmark built to isolate fluid reasoning and resist memorization, and beating the average human's 66%. The read everywhere is that frontier models have essentially cracked abstract reasoning.
Before you read on. Your call?
TRUE, BUT
92.5%
the score is real and ARC Prize verified, and on its own terms it is a genuine jump. Two things temper the headline. First, the 92.5% depends on Sol's most expensive reasoning mode; dial the effort down and ARC-AGI-2 collapses to 42.5%, and the ARC Prize's own cost-capped grand prize, which needs above 85% cheaply, is still unclaimed. Second, ARC-AGI is a series that keeps raising the bar, and the newest rung is a different kind of test. ARC-AGI-3 is an interactive, agentic benchmark, where a model must explore an unfamiliar environment, infer the goal, and plan, and there the same model at the same setting scores 7.78% while human testers solve 100%. That does not prove the ARC-AGI-2 result was hollow. It shows the fluid-reasoning progress does not yet extend to agentic novelty, and that every time this team builds a new test, today's models fail it until they catch up. Saturating a benchmark is not the same as acquiring the ability it was built to isolate.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →
Receipts
- Supports benchlm.ai:
leads the ARC-AGI-2 leaderboard with 92.5%
- Context benchlm.ai:
Average individual human performance is 66%
- Refutes arcprize.org:
Variant ARC-AGI-1 ARC-AGI-2 ARC-AGI-3 Max 96.5% 92.5% 7.78%
- Context arcprize.org:
High 97.0% 85.4% 2.15% Medium 92.5% 67.1% 1.07% Low 74.5% 42.5% 0.33%
- Context arxiv.org:
Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge
- Refutes arxiv.org:
Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%.
- Context arxiv.org:
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.