Single-Agent vs Multi-Agent Systems: When One Agent Is Enough
Single-agent vs multi-agent AI systems compared with evidence from Anthropic, Cognition and OpenAI: when multiple agents win, what they cost in tokens, and the patterns that actually work.
Contents
The most useful thing to know about multi-agent systems is that the teams who build the best ones disagree about when to use them, and the disagreement is informative.
In June 2025, one day apart, two agent-building teams published opposite-sounding posts. Cognition, the maker of Devin, titled theirs "Don't Build Multi-Agents" (Cognition, 2025). Anthropic published "How we built our multi-agent research system," reporting a 90.2% improvement over a single agent (Anthropic, 2025). Ten months later, Cognition published a follow-up describing the multi-agent patterns that do work (Cognition, 2026).
Read together, they agree more than they disagree. This article lays out where.
Definitions
- Single-agent system: one model with tools runs the loop from start to finish. It may call many tools, but one context holds every decision.
- Multi-agent system: work is split across several model instances, each with its own context, coordinated by a lead agent, by handoffs, or by a shared plan.
OpenAI's guide names the two common multi-agent shapes: a manager pattern, where a central agent calls specialists as tools, and a decentralised pattern, where agents hand control to one another (OpenAI guide). For the wider catalogue of architectures, see our AI agent architectures guide.
The case for one agent
Cognition's argument rests on two principles:
- "Share context, and share full agent traces, not just individual messages."
- "Actions carry implicit decisions, and conflicting decisions carry bad results."
Their example: split "build a Flappy Bird clone" into a background subagent and a bird subagent, and you may get a Super Mario-style background and a bird that looks and moves nothing like Flappy Bird's. Each subagent made reasonable choices the other could not see. Cognition's conclusion was that the "simplest way to follow the principles is to just use a single-threaded linear agent," which "will get you very far" (Cognition, 2025).
OpenAI's guidance points the same way: "maximize a single agent's capabilities first," because more agents add complexity and overhead, "so often a single agent with tools is sufficient" (OpenAI guide).
Single agents are also easier to debug, since one trace holds every decision, and cheaper, which the numbers below make clear.
The case for many
Anthropic's Research feature uses an orchestrator-worker design: a lead agent plans, then starts subagents that search in parallel, each in its own context window, and returns condensed findings. On Anthropic's internal research evaluation, a Claude Opus 4 lead with Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% (Anthropic, 2025).
Why it worked is the key finding. On the BrowseComp browsing benchmark, three factors explained 95% of performance variance, and token usage alone explained 80%. Multi-agent systems "work mainly because they help spend enough tokens to solve the problem": separate context windows add capacity for parallel reasoning that one window cannot hold.
Their example task shows the fit: finding all board members of the companies in the S&P 500 Information Technology sector. The multi-agent system split it into subagent tasks; the single agent "failed to find the answer with slow, sequential searches." Parallelism, with 3 to 5 subagents at once each using 3 or more tools in parallel, cut research time by up to 90% for complex queries.
The cost you pay
Anthropic is explicit about the downside:
| Interaction type | Tokens relative to a chat |
|---|---|
| Chat | 1× |
| Single agent | ~4× |
| Multi-agent system | ~15× |
Source: Anthropic, 2025
"For economic viability, multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance." Anthropic also names where they do not fit: domains "that require all agents to share the same context or involve many dependencies between agents." Most coding tasks, for example, "involve fewer truly parallelizable tasks than research."
There are coordination failures too. Early versions of Anthropic's system spawned 50 subagents for simple queries and searched endlessly for sources that did not exist. The fix was explicit effort rules in the prompt:
- Simple fact-finding: 1 agent, 3–10 tool calls
- Direct comparisons: 2–4 subagents, 10–15 calls each
- Complex research: more than 10 subagents with clearly divided responsibilities
Where the two camps agree
Cognition's 2026 follow-up reconciles the positions. Their original warnings "still hold today for parallel-writer swarms," but they found a class of patterns that work: "setups where multiple agents contribute intelligence to a task while writes stay single-threaded" (Cognition, 2026).
Their footnote makes the link explicit: both June 2025 posts "came to similar conclusions about the first area of applicability being in readonly agents." Anthropic's subagents search; one lead writes the answer. That is the same rule.
The patterns Cognition reports working in production:
- Clean-context reviewer. A second agent reviews the first agent's code without its history. On PRs Devin wrote itself, the reviewer catches an average of 2 bugs per PR, about 58% of them severe. It works best without shared context, because a short, fresh context avoids the degradation long contexts cause.
- "Smart friend." A cheaper primary model consults a stronger one on hard steps. This worked well between two frontier models; with a much weaker primary, it was "still an open problem."
- Manager and children. A manager splits a large task, child agents execute, the manager synthesises: "map-reduce-and-manage." Unstructured swarms of agents negotiating with one another are, in their words, "mostly a distraction."
Decision guide
| Your task… | Choose |
|---|---|
| Has steps that depend on each other's details (writing code, editing one document) | Single agent |
| Splits into independent read-only parts (research, search, data gathering) | Multi-agent, orchestrator-workers |
| Needs a quality check on the output | Single writer + clean-context reviewer |
| Overflows one context window | Subagents that return condensed summaries |
| Is cheap or high-volume | Single agent (15× tokens rarely pays) |
| Has a single agent picking the wrong tools among many similar ones | Split by domain, after first trying clearer tool descriptions |
The last row comes from OpenAI: some systems manage more than 15 distinct tools, while others struggle with fewer than 10 overlapping ones. Clarify the tools first, and split only if that fails.
Practical rules
- Start with one agent. Add agents to fix a measured problem, not a hypothetical one.
- Keep writes single-threaded. Parallelise reading, not deciding.
- Hand off full context or clear boundaries. Anthropic's subagents needed "an objective, an output format, guidance on the tools and sources to use, and clear task boundaries," otherwise they duplicated work.
- Write subagent outputs to files, not chat. Anthropic recommends subagents store work in external systems and pass back "lightweight references" to avoid a game of telephone.
- Budget explicitly. Encode effort rules and a token budget before you let a lead agent spawn workers.
- Trace everything. Multi-agent failures are emergent. Anthropic says "small changes to the lead agent can unpredictably change how subagents behave." You need per-agent traces to find out why. See AI agent observability.
How AGNT handles this
Disclosure: AGNT is our product. AGNT defaults to one agent per job, with an explicit tool list and credit limit. When a job needs more, a goal plans, runs steps and evaluates results against your criteria, and agents can delegate specific subtasks, including coding work to Claude Code or Codex. Every agent's run writes its own receipt, so a multi-agent run can be read step by step afterwards.
FAQ
Are multi-agent systems better than single agents?
For broad, parallelisable, read-heavy tasks like research, yes. Anthropic measured a 90.2% gain. For tasks with tightly linked decisions, such as coding, a single agent is usually more reliable and far cheaper.
How much more do multi-agent systems cost?
Anthropic measured about 15× the tokens of a chat interaction, against about 4× for a single agent.
What is the orchestrator-worker pattern?
A lead agent breaks the task down, delegates pieces to worker agents with their own context windows, and synthesises their results. It is the pattern behind Anthropic's Research feature.
What is context engineering?
Cognition's term for deciding what information each model call sees. It generalises prompt engineering to dynamic, multi-step systems, and it is the main reason multi-agent designs succeed or fail.
Sources
- Anthropic, How we built our multi-agent research system (June 2025)
- Cognition (Walden Yan), Don't Build Multi-Agents (June 2025)
- Cognition (Walden Yan), Multi-Agents: What's Actually Working (April 2026)
- OpenAI, A practical guide to building agents
- Chroma, Context Rot (July 2025)