Justin Bartak · AI Org · · 5 min read
Swarms Demo. Legions Ship.
TL;DR
A swarm of agents is emergent, anonymous, and spectacular in a demo. A legion is chartered, capped, attributed, and accountable, and it is what actually ships. I run my company as a legion: named seats, written charters, gates on every lane. The Romans beat bigger hordes with smaller numbers for four centuries. Formation beats enthusiasm.
Everyone selling multi-agent systems shows you a swarm: dozens of agents, self-organizing, intelligence emerging from the buzz. It is a spectacular demo. Then production arrives with its boring questions, who wrote this, who approved it, who owns the mistake, and the swarm has no answer, because having no answer is what emergence means. A legion answers all three.
Swarm or legion is not a size question. It is the question of whether output has an owner.
What is a swarm, honestly?
Many agents, coordinating by emergence. No fixed roles, no charters, no attribution, behavior nobody designed and nobody can fully reproduce. The appeal is real: emergence explores solution spaces that rigid structures never visit, and watching it work feels like watching intelligence.
The problem is that every property that makes a swarm mesmerizing makes it unshippable. Emergent behavior is unauditable by definition. Anonymous contribution means no artifact traces to a decision. And when the output is wrong, the swarm has no seat to hold accountable, only a vibe that misfired. I wrote about the extreme version of ungoverned replication in Hello, Agent Smith; the swarm is Smith's charming cousin.
What makes a legion different?
Structure, written down. The Roman legion was roughly 5,000 soldiers, organized into cohorts and centuries of 80 under named centurions, drilled until formation held under pressure, marching under a standard that made every unit's identity visible. Rome beat larger, braver, more enthusiastic hordes for four centuries with that structure. The horde had energy. The legion had an org chart.
My version, running Orbyt today: a 12-seat agent org where every seat holds a written charter naming its purpose and boundaries. Authority splits into 29 autonomous lanes and 40 gated ones, decided in advance by consequence. Output passes 103 mechanical guards. Outcomes land in a decision log at 40 entries. And the hard boundary: none of 12 seats can merge to the branch that deploys. The full governance layer is its own story; the point here is that every one of those properties is the opposite of emergence.
The horde has energy. The legion has a chart, and the chart is why it wins.
Does formation cap the scale?
No, and this is the part the swarm pitch gets wrong. Formation is what makes scale survivable.
My largest orchestrated workflow accumulated 298 agents in one run, holding to eight in parallel. One session spawned 830 agents across seven days, never more than nine at the same instant. One month of transcripts records 2,800 agent runs, 2,589 of them inside 95 orchestrated workflows. Those are swarm-sized numbers moving in formation: fan-out by design, caps by configuration, every unit's work signed and gated.
The legion's centuries and cohorts were exactly this: massive force, decomposed into units small enough to command. Orchestration is the century. The cap is the drill. The permission model is the chain of command.
What does the discipline cost?
Something specific, and I pay it in public.
Across the last 51 recorded seat runs: 43 completions and 30 discards at gates, an 84.3 percent completion rate. Roughly a third of produced work dies at the boundary rather than shipping unverified. That is the drill doing its job, and it is slower than letting the swarm buzz.
It also buys the thing the swarm can never sell: when something does go wrong, the failure has an address. The lesson lands in a corpus of 71 recorded incidents, and 47 of them are mechanized into permanent guards. A swarm's failure disperses back into the swarm. A legion's failure becomes doctrine.
When is a swarm the right tool?
Exploration, honestly bounded. Brainstorm fan-outs, search over hypotheses, adversarial review panels where diversity of attack matters more than attribution of finding. I run those shapes inside orchestrated workflows constantly. The distinction is that the swarm phase ends at a gate: everything it produces is treated as raw material for a chartered seat to verify, never as finished work.
A swarm inside a legion is reconnaissance. A swarm instead of a legion is a mob with API keys.
What to do Next
Ask the ownership question of every multi-agent system you run or are being sold. For any given artifact it produces: which agent made it, under what authority, and which human owns the outcome? If the answer is "it emerges," you have a demo, whatever the invoice says.
Then draft your first formation. Take your highest-volume agent use case and give it a charter, a cap, a gate, and a name. One chartered seat that ships beats forty anonymous agents that impress. Managing them starts with being able to address them.
And keep a small swarm, deliberately, for scouting, inside the walls. Rome did that too. The auxiliaries explored. The legion decided.
Enthusiasm disperses. Formation compounds. Build the legion, and let the swarm scout for it.
Related reading:
-
Hello, Agent Smith. what replication without formation becomes
-
The Machine. It Runs the Company. the charters, ladder, and kill switch behind the legion
-
391 Yeses and Not One No. the chain of command, measured
-
My C-Suite of Agents Named Themselves. the seats, introduced
-
AI Roadmaps Fail When They Ship Features the roadmap failures that make swarms inevitable
Originally published on orbytlabs.ai on Sep 6, 2026.
Frequently asked questions
What's the difference between a swarm of AI agents and a legion of AI agents?
A swarm is many agents coordinating by emergence, with no fixed roles, no charters, no attribution, and behavior nobody designed and nobody can fully reproduce. A legion is structure, written down, with charters naming purpose and boundaries. The real distinction is not size, it is whether the output produced has an owner, someone who can be named and held to it.
Why do multi-agent AI swarms fail in production?
Every property that makes a swarm mesmerizing in a demo makes it unshippable. Its behavior is unauditable by definition, and anonymous contribution means no artifact traces to a decision. When the output is wrong, the swarm has no seat to hold accountable, only a vibe that misfired. Production asks who wrote it, approved it, and owns the mistake, and emergence has no answer.
How is Orbyt's AI agent organization structured?
Orbyt is run as an agent org with 12 seats, and every seat holds a written charter naming its purpose and boundaries. Authority splits into 29 autonomous lanes and 40 gated lanes, decided in advance by consequence. Output passes 103 mechanical guards, outcomes land in a decision log at 40 entries, and none of the seats can merge to the branch that deploys.




