The trick: Lab Not Field
OpenAI and Anthropic are selling the same next step for AI agents: more of them.
Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.
parallel subagents are the next leap for AI agents. On August 13 Claude Code turned subagent forking on by default, OpenAI's Sol Ultra fans a task across up to 64 concurrent subagents, and the industry heuristic driving all of it is, in the words of the researchers who tested it, the belief that adding specialized agents will consistently improve results.
Before you read on. Your call?
TRUE, BUT
-70%
Google Research and academic collaborators ran the first controlled study large enough to answer it, 260 configurations across six agentic benchmarks and five architectures, holding tools, prompts, and compute fixed. The verdict is not that multi-agent is fake. It is that task structure decides. On a parallelizable benchmark like Finance-Agent, coordination delivered up to +81%. On a strictly sequential one like PlanCraft, every multi-agent variant they tested degraded performance by 39 to 70%. Their own summary: coordination improves performance on parallelizable tasks but degrades it on sequential ones. The open question for the two coding agents pushing subagents hardest is which pattern their default triggers: fanning out to gather context in parallel is the winning case, but splitting a dependent job like debugging or planning across coordinating agents is the losing one, and much of what developers hand these tools is the dependent, tool-heavy kind.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →
Receipts
- Supports code.claude.com:
Subagent forking is now on by default
- Supports digitalapplied.com:
Up to 64 concurrent subagents · unreviewed proof · 4 is the real default
- Context research.google:
believing that adding specialized agents will consistently improve results
- Refutes research.google:
parallelizable tasks like Finance-Agent (+81%) while degrading performance on sequential tasks like PlanCraft (-70%)
- Refutes research.google:
every multi-agent variant we tested degraded performance by 39-70%
- Context arxiv.org:
Across 260 configurations spanning six agentic benchmarks, five canonical architectures
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.