Subscribe

The trick: Lab Not Field

OpenAI and Anthropic are selling the same next step for AI agents: more of them.

Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.

Issue 1319 August 20266 receipts4 min

parallel subagents are the next leap for AI agents. On August 13 Claude Code turned subagent forking on by default, OpenAI's Sol Ultra fans a task across up to 64 concurrent subagents, and the industry heuristic driving all of it is, in the words of the researchers who tested it, the belief that adding specialized agents will consistently improve results.

Before you read on. Your call?

Google Research and academic collaborators ran the first controlled study large enough to answer it, 260 configurations across six agentic benchmarks and five architectures, holding tools, prompts, and compute fixed. The verdict is not that multi-agent is fake. It is that task structure decides. On a parallelizable benchmark like Finance-Agent, coordination delivered up to +81%. On a strictly sequential one like PlanCraft, every multi-agent variant they tested degraded performance by 39 to 70%. Their own summary: coordination improves performance on parallelizable tasks but degrades it on sequential ones. The open question for the two coding agents pushing subagents hardest is which pattern their default triggers: fanning out to gather context in parallel is the winning case, but splitting a dependent job like debugging or planning across coordinating agents is the losing one, and much of what developers hand these tools is the dependent, tool-heavy kind.

-70%WORST MULTI-AGENT DEGRADATION ON A STRICTLY SEQUENTIAL TASK
39-70%RANGE BY WHICH EVERY MULTI-AGENT VARIANT DEGRADED SEQUENTIAL REASONING
+81%WHERE MORE AGENTS DO HELP: A PARALLELIZABLE TASK
64CONCURRENT SUBAGENTS OPENAI'S SOL ULTRA CAN FAN A PROBLEM ACROSS

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Both big coding agents now push parallel subagents as the win, but the controlled Google Research study of 260 configs found more agents help parallelizable tasks by up to 81% and hurt sequential ones like planning by up to 70%. Task structure decides, not agent count, so a default that fans out for independent subtasks can help while one that splits a dependent chain is a tax.”

Receipts

  1. Supports code.claude.com: Subagent forking is now on by default
  2. Supports digitalapplied.com: Up to 64 concurrent subagents · unreviewed proof · 4 is the real default
  3. Context research.google: believing that adding specialized agents will consistently improve results
  4. Refutes research.google: parallelizable tasks like Finance-Agent (+81%) while degrading performance on sequential tasks like PlanCraft (-70%)
  5. Refutes research.google: every multi-agent variant we tested degraded performance by 39-70%
  6. Context arxiv.org: Across 260 configurations spanning six agentic benchmarks, five canonical architectures

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.