Oxford philosopher Toby Ord's math says parallel agents buy speed but pay a coordination tax where workers duplicate each other's work. A 10 agent AI swarm delivers roughly 3–5x useful output, not 10x, and the gap widens with scale.
Four AI agents running in parallel can finish a job in roughly half the wall-clock time of one agent, while burning about twice the tokens to do it. "Tokens" here are the small chunks of text a model reads and writes; a swarm spends more of them because each agent is doing its own pass. That is the simple picture. The harder math, drawn out by the independent researcher Toby Ord in an analysis on his blog, is what happens as you keep adding agents: a 10-agent swarm is closer to three to five times as useful as a single agent, not ten times, and the gap widens further with scale.
The mechanism has a name. Ord calls it "stepping on toes": as more agents work the same problem in parallel, they begin duplicating effort, contradicting each other's intermediate steps, or stalling on shared context. The tax compounds. Ord models swarm useful work as roughly 10^λ for λ between 0.48 and 0.68, a range derived from three benchmarks—Terminal-Bench (λ ≈ 0.48), SEC-Bench Pro (λ ≈ 0.57), and BrowseComp (λ ≈ 0.68). A 10x increase in agents yields only about 3x to 5x in delivered output; a 100x increase drops the multiplier to roughly 10x to 30x. That is a sharp curve, and it sits in the opposite direction from how the "just spin up more agents" pitch tends to land.
That math reframes what an "AI swarm" actually is. Capability in modern AI has long had two scaling axes: bigger models, and more inference compute (the chips and time spent running a trained model) applied to a single model at test time. A swarm of cooperating agents is a third axis, and the one with the steepest diminishing returns. Parallel runs buy wall-clock time, which matters when a task is genuinely parallelizable and latency-bound. They cost more tokens for the same problem, and they pay a coordination overhead that grows faster than the headcount.
Ord does not treat the result as a free lunch. In his own framing, swarm scaling adds a third knob without an obvious ceiling, and is one more vector that could feed a recursive self-improvement loop, where AI systems improve the AI systems that improve themselves. His read is that the math leaves a recursive-self-improvement-driven intelligence explosion somewhat more plausible than it would be without swarms, not less. That is a caveat, not a verdict, but it belongs in the picture rather than buried at the bottom.
The practical reading is sharper than either the doom or the cheerleading version of the swarm story. When a workload is embarrassingly parallel and latency-bound, where the work splits cleanly into independent pieces and the user is waiting on the clock, running more agents is a real speed lever, and the token cost is the price of admission. When a workload is sequential or tightly coupled, where each step depends on the previous result, the coordination tax eats the savings quickly, and the right move is to spend the budget on a single stronger model or more inference compute instead. The third scaling axis is real, and the trade-off is named: a swarm buys time, not capability, and only when the work is actually parallel.
The next signal is the price sheet. When model labs start publishing per-token pricing for "swarm" or "agent-fleet" products in their APIs, the structure of those SKUs will signal whether vendors expect teams to use this third axis as a default tool or as a niche escape hatch for latency-bound work.