AI Agent Orchestration: Designing and Implementing Multi-Agent Systems
One failing agent can erase the work of the ones that succeeded. Four orchestration patterns, plus the boundaries that break in production, with runnable fixes.
Three agents run in parallel. One of them hits a rate limit. And just like that, the output from the two that already finished disappears along with it. That was the first floorboard I put a foot through when I tried to move a multi-agent setup into production.
The cause had nothing to do with grand architectural choices. It was a plain loop over future.result(). The first exception propagates immediately, and the function exits before it ever collects what the other agents had already returned.
The hard part of orchestration lives in those boundary implementations far more often than in pattern selection. So this piece lays out the four core patterns, then follows an orchestrator into the places where it actually breaks — with code you can run.
One note: this article has been revised since publication. The original implementation had three holes in it — partial failure handling, extracting JSON from LLM responses, and timestamps. Those sections now carry corrected code, along with an explanation of why the originals broke.
Why Multi-Agent Systems?
Single agents are highly efficient for well-scoped tasks. But real-world business processes rarely fit that mold.
Limits of Single Agents
Context window constraints: Even the latest LLMs have limits on how much information they can process at once. Analyzing large documents or handling multi-step complex tasks quickly runs into this ceiling.
Lack of specialization: Asking one agent to handle everything leads to bloated prompts and declining output quality — the equivalent of expecting one person to be both a CPA and a legal expert.
No parallelism: Single agents are inherently sequential. Even when tasks A and B are entirely independent, one has to finish before the other can start.
Error propagation risk: When a single agent fails, the entire workflow stops. With separated agents, partial failures are far less likely to cascade.
What Multi-Agent Systems Solve
Multi-agent systems address these issues directly. Each agent has a clearly defined role and access only to the tools that role requires. Communication between agents follows a structured protocol that enables parallel execution. And partial failures no longer bring the whole system down.
Four Core Orchestration Patterns
Pattern 1: Centralized Orchestrator
The most common pattern. A central orchestrator makes all decisions and dispatches instructions to sub-agents.
User
↓
Orchestrator (central command)
├── Instruction → Sub-Agent A
├── Instruction → Sub-Agent B
└── Instruction → Sub-Agent C
↑
Aggregates results and returns to user
Advantages: Entire system state is managed in one place, making debugging straightforward. Task dependencies are explicitly controlled.
Disadvantages: The orchestrator itself becomes a single point of failure. Its context can grow unwieldy over time.
Best for: Workflows with complex inter-task dependencies where strict execution order matters.
Pattern 2: Distributed Peer-to-Peer
Agents communicate directly with one another — no central command.
Agent A ←→ Agent B
↕ ↕
Agent C ←→ Agent D
Advantages: No single point of failure. Each agent can scale independently.
Disadvantages: Overall system state is harder to observe. Risk of deadlocks or infinite loops.
Best for: Clearly delineated, highly independent agent roles. Peer review or mutual verification use cases.
Pattern 3: Hierarchical Multi-Level
A top-level orchestrator manages multiple intermediate managers, each of which oversees leaf agents.
Top Orchestrator
├── Manager A
│ ├── Worker A1
│ └── Worker A2
└── Manager B
├── Worker B1
└── Worker B2
Advantages: Scalable to large systems. Clear separation of responsibilities at each level.
Disadvantages: Increased latency. Communication overhead between layers.
Best for: Large-scale workflows integrating multiple independent subsystems.
Pattern 4: Dynamic Agent Spawning
The orchestrator creates and destroys agents on the fly, based on what each task actually requires.
def dynamic_orchestrator(task: str) -> str: """Dynamically spawn agents based on task analysis""" task_analysis = analyze_task(task) required_agents = task_analysis["agents_needed"] active_agents = {} for agent_spec in required_agents: active_agents[agent_spec["id"]] = create_agent( role=agent_spec["role"], tools=agent_spec["tools"], system_prompt=agent_spec["prompt"] ) results = execute_with_agents(task, active_agents) for agent in active_agents.values(): agent.cleanup() return results
Advantages: Efficient resource use. Agent configuration is optimized per task.
Best for: Highly varied task types where the required agents can't be determined in advance.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦Key architectural patterns for multi-agent systems and how to choose the right one
✦Implementing orchestrators, sub-agent roles, communication protocols, and state management
✦Practical approaches to scaling, reliability, and cost challenges in production multi-agent deployments
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Building an Orchestrator: Implementation Deep Dive
Task Decomposition Engine
The core of any orchestrator is its ability to break complex tasks into executable sub-tasks.
First, a shared utility. LLMs often wrap their answer in a preamble like "Sure, here's the plan" or enclose it in a code fence, and json.loads() falls over the moment they do — a fenced response passed straight to json.loads() returns JSONDecodeError: Expecting value.
import jsonimport redef extract_json(text: str) -> dict: """Pull the first JSON object/array out of an LLM response.""" fence = chr(96) * 3 # a code fence: three backticks body = text.strip() fenced = re.search(fence + r"(?:json)?\s*(.+?)\s*" + fence, body, re.S) if fenced: body = fenced.group(1).strip() starts = [i for i in (body.find("{"), body.find("[")) if i != -1] if not starts: raise ValueError(f"JSON not found in model output: {text[:200]}") # raw_decode tolerates prose trailing after the JSON obj, _ = json.JSONDecoder().raw_decode(body[min(starts):]) return obj
Fenced output, bare JSON, and JSON with commentary tacked on the end all pass through this single function. Structured output (tool use) is the sturdier choice when it's available, but while you're still swapping models around during prototyping, having this catch-all in place visibly cuts down on breakage.
from anthropic import Anthropicimport jsonclient = Anthropic()def decompose_task(task: str, available_agents: list[dict]) -> list[dict]: """ Use an LLM to decompose a task and assign sub-tasks to agents """ decompose_prompt = f""" You are a task decomposition expert. Break the following task into sub-tasks that can be executed by the available agents. Main task: {task} Available agents: {json.dumps(available_agents, indent=2)} Respond in this JSON format: {{ "subtasks": [ {{ "id": "task_1", "description": "Sub-task description", "assigned_agent": "agent_id", "depends_on": [], "can_parallel": true }} ], "execution_order": [["task_1", "task_2"], ["task_3"]] }} execution_order is a 2D array where each inner array contains tasks that can run in parallel. """ response = client.messages.create( model="claude-opus-5", max_tokens=4096, messages=[{"role": "user", "content": decompose_prompt}] ) return extract_json(response.content[0].text)import concurrent.futuresdef execute_task_plan(plan: dict, agents: dict) -> dict: """Execute a task plan sequentially and in parallel""" results: dict = {} failures: dict = {} for parallel_group in plan["execution_order"]: if len(parallel_group) == 1: task_id = parallel_group[0] subtask = next(t for t in plan["subtasks"] if t["id"] == task_id) agent = agents[subtask["assigned_agent"]] context = {dep: results[dep] for dep in subtask["depends_on"]} results[task_id] = agent.execute(subtask["description"], context) else: # Parallel execution — one failure must not erase the others with concurrent.futures.ThreadPoolExecutor(max_workers=8) as executor: futures = {} for task_id in parallel_group: subtask = next(t for t in plan["subtasks"] if t["id"] == task_id) agent = agents[subtask["assigned_agent"]] context = {dep: results[dep] for dep in subtask["depends_on"]} futures[executor.submit( agent.execute, subtask["description"], context )] = task_id for future in concurrent.futures.as_completed(futures): task_id = futures[future] try: results[task_id] = future.result() except Exception as exc: failures[task_id] = f"{type(exc).__name__}: {exc}" return {"results": results, "failures": failures}
This is the floorboard from the opening. The original read for task_id, future in futures.items(): results[task_id] = future.result(). With that shape, the moment the first future you touch happens to hold an exception, the whole of execute_task_plan unwinds. Two of your three agents may have succeeded — you still walk away with nothing but their token bill.
Collecting with as_completed and isolating failures in a try/except keeps the successful work. Running this locally with one of three agents forced into a rate limit returns {'results': {'c': 'c:done', 'a': 'a:done'}, 'failures': {'b': 'RuntimeError: rate limit'}} — two successes and one failure, cleanly separated.
Note that the return shape changes. Callers now check whether failures is empty and decide whether to retry or proceed with a gap. The point isn't to swallow partial failure; it's to put it somewhere you can see it.
Shared Workspace for Agent State
Multi-agent systems need a mechanism for sharing state across agents.
In distributed systems, observability isn't optional. Trace each agent call, correlate agent IDs with workflow IDs, and use structured logging so you can reconstruct what happened when something goes wrong.
Security and Governance
Agent Identity and Authentication
Assign each agent a unique identity and verify that instructions originate from trusted sources. This prevents impersonation between agents.
Permission Scoping
Each agent should only have access to the tools its role requires. Separate read-only agents from write-access agents explicitly.
Audit Logging
Record every inter-agent communication and external tool call. This enables root-cause analysis and supports compliance requirements.
The Time I Dismantled an Orchestration
I took one of these setups apart while working on metadata for a wallpaper app. Each image needed three things done to it — category classification, description generation, and tagging — so I assigned all three to separate agents. The roles were distinct, so surely they should be separate. That was the reasoning.
What actually grew was the number of API calls and the number of failure combinations. Classification fails but the description gets written anyway. Tags stay stale from the previous run. Untangling those half-finished states turned into logic more complicated than the pipeline itself.
Today only the classification step remains an agent. Descriptions and tags are built deterministically from the classification result. The only thing that genuinely requires judgment is "what is this image," and everything downstream of that fits in a lookup table.
Agents earn their keep where the input is ambiguous and a judgment call is required. Turn a step whose output is fully determined by its input into an agent, and all you've added is nondeterminism and new failure modes. Before you pick a pattern, walk through each step asking whether it truly needs judgment — the agent count usually comes out lower than you expected.
Working solo, the person who gets paged when something falls over is me. So I've come to treat "add one more agent" as identical to "add one more thing to monitor and one more way to fail."
Principles That Stuck
These are the judgment calls that kept paying off across the implementations and the failures above.
Start simple: Don't reach for multi-agent complexity until you've confirmed a single agent can't solve the problem. Add complexity only when you have a clear reason.
Clarify roles: Define each agent's scope precisely and minimize overlap. The "one agent, one responsibility" principle improves maintainability considerably.
Build observability in from day one: Logs, metrics, and tracing are design decisions, not afterthoughts.
Design for failure: Agents will fail. Implement retries, fallbacks, and circuit breakers from the start so partial failures don't bring the whole system down.
Track costs actively: Monitor every agent call and API invocation. Mix model tiers and use prompt caching aggressively.
If you want one concrete next step: open the parallel execution path in whatever workflow you're running right now and read the single line where future.result() gets called. If exceptions propagate on the spot, you're probably losing successful work. Swapping in as_completed and a try/except takes a dozen lines, and the effect is immediately measurable.
I had the same hole sitting in this article's own code. Sometimes it pays to doubt a few lines at the boundary before reaching for the architecture diagram. Thanks for reading.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.