AgentKit 2.0 Production Design Patterns — Strategic Architecture for 16 Specialized Agents
How to compose AgentKit 2.0's 16 specialized agents on a real production project: five orchestration patterns, which four agents earn a permanent seat, AGENTS.md as team standards, and how to split work across a 1M token context with Gemini 3.1 Pro.
The Monday morning after I ran all sixteen agents at once
The first weekend I had AgentKit 2.0 installed, I launched all sixteen agents at the same time. Frontend, backend, security audit — run them in parallel and surely the skeleton of the project would be standing by morning.
What was waiting on Monday was four mutually contradictory API specifications and a repository where nobody could say which branch was authoritative.
Working as an indie developer, being short-handed is the normal condition, which is exactly why I never questioned the instinct that parallelism equals speed. In agent workflows that instinct is only half right. Parallelism pays off when the decomposition is correct. Get the decomposition wrong and raising concurrency simply inflates the integration cost first.
What follows is the architecture I rebuilt after that weekend. It doesn't begin with a feature tour of the sixteen agents. It begins with the wiring: who receives what, and in what order.
Chapter 1: The Full Picture — Understanding AgentKit 2.0's 16 Agents
Agent Classification Framework
AgentKit 2.0's 16 agents naturally cluster into six functional domains:
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦Understanding 16 specialized agents and designing optimal team compositions for different project scales
✦Five production-ready orchestration patterns: hierarchical delegation, pipeline, fan-out, and more
✦Team-unified rules via AGENTS.md and large-scale development strategies leveraging 1M token context
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Looking at the classification table, the temptation is to deploy all sixteen. In practice I keep four running continuously. The other twelve get woken up only when a specific situation calls for them.
The four that stay resident
Project Coordinator — the only one always on. Its real job is deciding whether the other three need to be called at all
Backend Architect — once the schema and endpoint shapes are settled, downstream rework drops sharply
Test Engineer — agent-generated code tends to be "working but brittle," so having tests written first ends up being faster
Security Auditor — most of the rejections I've seen in App Store review trace back to skipping this step
The twelve I call situationally
Design System Manager and Cloud Architect only pay for themselves once the surface area crosses a certain size. On an app with fewer than ten screens, asking for a full design-token system produced token definitions larger than the implementation itself, and every subsequent change got heavier rather than lighter.
I decide residency on three questions:
Will I reference this output three or more times? If it's consulted once or twice, handle it inline instead of standing up a dedicated agent
Does the output feed another agent? Terminal artifacts such as reports can be generated on demand
Can a human articulate the spec? Delegating a domain you can't articulate returns work you can't review
Running all sixteen means personally carrying sixteen reviewers' worth of accountability. The real concurrency ceiling isn't your plan's parallel-execution limit — it's how many outputs you can actually review. Adopting that framing changed how the tool behaves more than any configuration change did.
I set concurrency from last night's diff size
I wrote "how many outputs you can review," but the count itself turned out to be a poor unit. A branch that changes twelve lines and a branch that rewrites nine hundred were both landing in my queue as "one." These days I measure the lines I'll have to read in the morning, not the branches.
Every branch that ran overnight gets inventoried first thing with this script.
#!/usr/bin/env bash# review-load.sh - measure the morning reading load per agent branch# usage: ./review-load.sh origin/mainset -euo pipefailBASE="${1:-origin/main}"BUDGET_LINES="${REVIEW_BUDGET_LINES:-600}" # what I can actually read before lunchtotal=0git fetch --quiet --allwhile read -r branch; do [ -z "$branch" ] && continue # lockfiles and generated output are not review surface lines=$(git diff --numstat "$BASE...$branch" -- \ ':!*.lock' ':!*-lock.json' ':!*.snap' ':!dist/**' ':!**/generated/**' \ | awk '{ added += $1; removed += $2 } END { print added + removed + 0 }') files=$(git diff --name-only "$BASE...$branch" | wc -l | tr -d ' ') printf '%-44s %6s lines %4s files\n' "${branch#origin/}" "$lines" "$files" total=$(( total + lines ))done < <(git for-each-ref --format='%(refname:short)' 'refs/remotes/origin/agent/*')printf '%-44s %6s lines (budget %s)\n' '--- TOTAL ---' "$total" "$BUDGET_LINES"if [ "$total" -gt "$BUDGET_LINES" ]; then echo "over budget: wake fewer agents, or push merges to tomorrow" exit 1fiexit 0
The agent/* branch naming exists for exactly one line of that script — the git for-each-ref filter. Without a naming rule decided up front, you end up picking the branches to inventory by eye every morning, and that selection is itself review time.
On mornings when it exits non-zero, I don't dispatch anything new. Starting fresh work on top of unread lines splits every later integration conflict into "yesterday's" and "today's," and the cause stops being legible.
Mapping line counts to actual minutes took a while to calibrate. From my own notes, it lands roughly here.
Changed lines in a branch
Time to read it
What I do
Under 80
Under 10 min
Read it through and merge the same day
80 - 300
20 - 40 min
Read the test diff first, then follow the implementation
300 - 900
An hour or more
Don't merge; send it back to be split
Over 900
I can't finish it
Treat it as a sign the task was scoped wrong and return the whole thing
That review quality degrades as the diff grows is obvious in hindsight, and I missed it for a long time. The code from the period when I was calling a nine-hundred-line branch "reviewed" after an hour still reads rough to me.
The plan sets your concurrency limit; your morning sets your parallelism. Since drawing that line I've stopped spending any time wondering whether to buy a bigger tier.
Chapter 2: Five Battle-Tested Orchestration Patterns
Pattern 1: Hierarchical Delegation
When to use: Large projects (6+ months) requiring staged responsibility distribution.
Explicit context handoff — Parent agents must provide children with:
The full project picture (where their work fits)
Interface boundaries (who collaborates with whom)
Constraints (tech stack, timeline, budget)
Success criteria (definition of "done")
Manager View inbox tracking — Centralize completion reports to optimize overall schedule and detect bottlenecks early.
Clear escalation paths — When blockers emerge, teams report upward through one level; the coordinator rebalances priorities.
How I use it: When payment work stalls, the only thing I look for in the Backend Architect's report is a line like "waiting on the Security Auditor's decision about the auth scheme." A stalled report with no stated reason is usually not a wait at all — it's the opening move of rework. What hierarchical delegation actually bought me wasn't depth. It was the reason for the stall surfacing one level up without me asking.
Pattern 2: Pipeline Pattern
When to use: Frequent feature releases (monthly or more).
Spec → Code Gen → Code Review → Test → Docs → Release
↓ ↓ ↓ ↓ ↓
Writer Builders Reviewers Test Docs
(Frontend Team Expert
+Backend)
Implementation essentials:
Sharp input/output contracts — The spec stage must be 100% complete (all endpoints, request/response types, error codes). The code generation stage never blocks waiting for clarification.
Identify parallelizable stages — Frontend UI and backend test suite creation can run in parallel. Visualize the critical path and compress it.
Quality gates between stages — Code passes linting and type checks before moving to testing; tests pass coverage thresholds before performance review. Automate, don't manually approve.
How I use it: What the pipeline shortened for me wasn't code generation time. It was the number of round trips on the spec. Hand a vague contract to the next stage and the downstream agent fills the gap by guessing — then peeling that guess back out lands at the very end. Now I don't wake the next stage until the spec stage has written out the shape of the error responses too.
Pattern 3: Fan-out/Aggregate
When to use: Independent module development or multi-environment deployments.
Set timeouts — Each parallel task has a maximum wait (default 30 min, critical tasks 60 min). Detect hung agents early.
Graceful partial failure — If Dev deployment succeeds but Staging fails, retry only Staging. Full re-execution wastes time.
Result aggregation — Manager View shows "3/3 passed" or "2/3 passed, Staging error: DNS config" for quick diagnosis.
How I use it: The pre-release check for my wallpaper app settled into exactly this shape. Duplicate detection, missing metadata, untranslated strings — three checks in parallel, and only the ones that fail get rebuilt. Back when a single failure meant redoing the whole batch, replacing one image could cost me an entire evening.
Pattern 4: Spec-Driven Development
When to use: Requirements are volatile and multiple refinements are expected.
Manager: "Define requirements for feature X"
↓
Backend Architect + Frontend Builder (co-author spec v1.0)
↓ (approved by Manager View)
Project Coordinator (schedule + resources)
↓
All development agents (implement spec v1.0)
↓ (6 weeks later)
Manager: "Customer wants module Y to work differently"
↓
Update spec → v1.1 → impact analysis → schedule adjustment → re-deploy
The specification becomes the single source of truth—not code, not Slack messages, not the manager's memory.
Implementation essentials:
Version control your spec — spec-v1.0.md → spec-v1.1.md (added security requirement) → spec-v2.0.md (API redesign). Always cite version in agent instructions.
Prevent spec drift — If code diverges from spec, stop implementation until they're aligned. Create a new spec version and notify all agents.
Formalize change requests — "I want to change X" → update spec → analyze impact → re-estimate schedule → notify agents. No ad-hoc modifications.
How I use it: You can't stop requirements from moving. What you can stop is a moved requirement that nobody wrote down while implementation keeps going. Write the change into the spec the day you think of it, have the agents read the diff, then resume — whether I hold that order or not is what visibly changes how much rework arrives in the second half.
Pattern 5: Quality Assurance Pipeline
When to use: Production downtime carries real cost (finance, healthcare, e-commerce).
Automate everything possible — If you're manually checking something twice, automate it. Eliminate human error.
Run gates in parallel — Code Quality and Test gates are independent; Performance and Security are independent. Compress 10-day sequential to 6-day parallel.
Handle failures decisively — "Performance baseline not met" → identify bottleneck → optimize → re-test. Don't defer to the next release.
Chapter 3: AGENTS.md — Your Team's DNA
The single most important artifact in AgentKit 2.0 production deployments is AGENTS.md: a document encoding every technical decision, standard, and safety rail that all 16 agents follow.
Essential sections in AGENTS.md
# AGENTS.md — Team Development Standards## Tech Stack Decisions- Frontend: React 18 + TypeScript 5.x + Tailwind CSS 3.x- Backend: Node.js 20 LTS + Express 4.x- Database: PostgreSQL 15 (primary), Redis 7.x (cache)- Testing: Vitest (frontend), Jest (backend)- Deploy: Docker + Kubernetes- Monitoring: Prometheus + Grafana## Code Quality Rules- TypeScript strict mode enabled- ESLint + Prettier mandatory- Test coverage minimum 80%- Cyclomatic complexity < 10- Line length ≤ 100 characters## Security Requirements- All env vars in .env.example (no hardcoding)- Auth: OAuth 2.0 Bearer tokens- DB: Prepared statements only (prevent SQL injection)- CORS: Explicit whitelist- Dependencies: Weekly npm audit, auto-patch low severity## Inter-Agent Communication- JSON format for all task specs- Dependencies as DAG (Directed Acyclic Graph)- Deliverables via Git (commits, PRs, tags)## Deployment Stages- Three environments: dev, staging, production- Canary deployment: 5% → 25% → 100%- Automatic rollback on error rate > 1%
Why does AgentKit 2.0 use Gemini 3.1 Pro? Its 1 million token context window (≈750,000 words). Every vendor's ceiling moves with each generation, so rather than memorizing multiples against someone else's model, I'd measure the number that actually matters to you: how many KB you can hand over at once before quality falls off.
Three-stage context utilization
Stage 1: Project-wide understanding on first interaction
1M token allocation:
- AGENTS.md: 15 KB
- Full specification: 200 KB
- Critical codebase (curated): 400 KB
- Dependency documentation: 100 KB
- Available for instructions: ~200 KB
When the Backend Architect receives its first assignment, it gets:
The entire tech stack and AGENTS.md
All API specs (no waiting for clarification)
6 months of Git history (understands past decisions)
Existing related code (avoids reinventing)
Result: No "wait, which endpoint was that?" delays.
Manager: "Here's last year's full codebase (800 MB).
Implement feature X using this."
→ 3-hour processing, $10+ cost, no better result.
✓ Right:
Curate 500 KB:
- AGENTS.md
- This month's changes
- Relevant existing code
- Past 3 months' commits
→ 1-min processing, $0.30, full context awareness.
Chapter 6: Composing a four-month build, start to finish
What follows is not a case report. It's a planning sketch: assume a team work-management SaaS with a four-month window, and lay the patterns above into an actual sequence. Treat every number as a slot to replace with your own measurements.
Assumptions I'm working from
Duration: 4 months
Humans: 2, both in the Project Coordinator role
Target: Web + iOS + Android
Phase 1: Discovery & Design (Weeks 1–2)
Manager View instructs:
"Based on business plan, define:
1. Core features (task creation, sharing, notifications, analytics)
2. Tech stack (React 18, Node 20, Postgres, Kubernetes)
3. Deployment model (SaaS + on-prem)
Agent assignments:
- Backend Architect: API structure, data model
- Frontend Builder: UI wireframes, interaction flow
- Database Engineer: Normalization, indices
- Security Auditor: Auth strategy, GDPR readiness
- Cloud Architect: AWS design (ECS, RDS, CDN)
- DevOps Engineer: CI/CD skeleton
"
Deliverables:
- Specification v1.0 (150 pages, final)
- AGENTS.md (v1.0)
- Architecture Decision Records (3 key decisions)
Phase 2: MVP Build (Weeks 3–10)
Hybrid approach: pipeline + fan-out
Weeks 3-4 (Parallel):
├→ Backend: Implement all APIs (spec-driven)
├→ Frontend: Build UI component library
└→ Infra: Docker dev environment
Weeks 5-6 (Parallel):
├→ Frontend: Integrate with API, build flows
├→ Backend: WebSocket notifications, real-time
└→ Test: Write E2E test scenarios
Weeks 7-8 (Parallel):
├→ Integration: Front+Back combined testing
├→ Security: OWASP review, dependency audit
└→ Performance: Lighthouse, API latency profiling
Weeks 9-10 (Parallel):
├→ Mobile: React Native implementation
├→ Docs: Auto-generate API specs
└→ Staging: Deploy and monitor
The sequencing is the point here, not the calendar: nothing in a later block may depend on an unfinished item in its own block.
Instead of quoting someone else's results, here are the four numbers I watch every week on my own sites.
Time lost to integration — how long it took to reconcile artifacts built in parallel. If this grows the week you raise concurrency, your decomposition is too coarse
Review queue depth — outputs that are finished but that I haven't read yet. Past three, I stop waking new agents
Times you returned to the spec — rewriting the spec mid-implementation isn't the problem. The times you didn't rewrite it are what come due later
Critical bugs in the first month after launch — the slowest and most honest signal that the QA pipeline is doing its job
Chapter 7: Cost & Parallelization Trade-offs
Each plan comes with a ceiling on how many agents you can run at once. Both the ceilings and the prices move with plan revisions, so what follows is the shape of the decision rather than a price list — check the official pricing page for current numbers.
Tier
Rough concurrency
Who it fits
Free
A couple of agents
Prototypes and one-off features, on the assumption you're sequencing rather than parallelizing
Individual paid
Around five
Steady solo development. This is where I've stayed
Higher tiers
Effectively unlimited
Teams that can split review across several people
Run the math on your own hours, not someone else's spreadsheet
The formula is simple. What makes it feel uncertain is borrowing the inputs from a case study that isn't yours.
Monthly decision input =
tool subscription
+ (hours you spend reviewing and coordinating x your hourly value)
- (extra days if you dropped back to sequential x what those days cost you)
The third term is the vague one, and I'd rather leave it vague than fake precision. As an indie developer I can't put a defensible figure on "shipping three weeks later," so I substitute something I can count: the articles I could have written and the bugs I could have fixed in those three weeks.
Whether to move up a tier, I decide from the script earlier in this article — specifically, how many mornings a week it exits non-zero. Most weeks the answer wasn't that the tier was too small. It was that I hadn't finished reading.
Chapter 8: Pre-Launch Checklist
Before release to production
[ ] AGENTS.md v1.0 complete, all 16 agents have read access
[ ] Specification finalized (v1.0 or higher), both teams synchronized
[ ] AGENTS.md reviewed for drift (tech updates, new best practices)
[ ] Dependency documentation refreshed
[ ] Past month incidents analyzed → AGENTS.md improvements
[ ] New CVEs evaluated, security guidelines updated
[ ] Cost analysis: tool spend vs. the hours you personally spent reviewing
Conclusion
All five patterns here — hierarchical delegation, pipeline, fan-out/aggregate, spec-driven, quality assurance — rest on the same precondition: that the decomposition is correct. Which means that when progress stalls, the thing to question is neither your concurrency setting nor your plan tier. It's the granularity of the split.
The concrete next step I'd suggest is writing twenty lines of AGENTS.md. Three sections only: tech stack, code quality rules, prohibited patterns. That alone visibly reduces contradictions between agents. A comprehensive AGENTS.md is something you grow through operation, not something you finish on day one.
My own first AGENTS.md was under twenty lines. It's several times that now, and every line added came from an incident that actually happened. It isn't a document assembled by filling in a template — it's a record of things that went wrong, which is precisely why it gets followed.
Thanks for reading this far. I hope it gives you a usable starting point.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.