Antigravity × A2A Protocol: Complete Implementation Guide for Scalable Multi-Agent Systems
A complete guide to implementing Google's A2A (Agent-to-Agent) protocol with Antigravity. From agent card design to task delegation, streaming communication, and production deployment on Cloudflare Workers.
Setup and context — How A2A Is Reshaping Agent Architecture
The day I stood up my first agent, I spent half an hour hammering curl trying to work out why my own client kept getting a 404. It wasn't auth, and it wasn't CORS. The article I'd copied the path from had been written before the spec settled, and the card was simply not where I was looking for it. What had fallen behind wasn't the code — it was me.
The A2A (Agent-to-Agent) protocol is an open standard for letting AI agents communicate and delegate tasks to one another. It was released in April 2025 and is now stewarded by the Agentic AI Foundation under the Linux Foundation. The spec has reached v1.0, and operation names and Agent Card structure both changed along the way. This is an area where guides with working code go stale fast, so I've put links to the primary sources I actually relied on inline throughout.
Before A2A, building multi-agent systems meant inventing a proprietary communication layer for each project. A2A changed that by providing a standard HTTP/SSE-based interface so agents built with different frameworks and languages can interoperate out of the box.
This guide walks you through building an A2A-compliant agent system with Antigravity, from first principles to production deployment. We cover agent card design, task delegation patterns, streaming, authentication, error handling, and deploying to Cloudflare Workers — everything you need to ship a robust multi-agent system.
Before writing any code, let's nail down the four building blocks of A2A.
Agent Card
The Agent Card is the agent's "business card" — a JSON manifest published at /.well-known/agent-card.json. It declares what the agent can do (skills), how to reach it (endpoint, auth schemes), and what input/output formats it supports.
One note on that path, since it cost me an afternoon. Early samples and older write-ups often show /.well-known/agent.json, but the standard path defined in the A2A Agent Discovery docs is /.well-known/agent-card.json, following the well-known URI conventions of RFC 8615. I pasted code from a transitional article and couldn't work out why my client kept getting a 404. If you need to stay reachable to peers mid-migration, serve the same card at both paths and retire the old one once they've moved.
A Task is the unit of work exchanged between agents. Each task has a unique ID and follows a lifecycle: submitted → working → completed (or failed). Tasks carry Messages (the conversation history) and Artifacts (the outputs).
Artifact
An Artifact is the deliverable produced by a task. It can contain multiple parts — text, files, or structured data — and can be streamed incrementally as the agent works.
Push Notifications
For long-running tasks, agents can notify clients via webhook when work is complete. This lets clients avoid holding open a persistent connection, making it ideal for batch workloads.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦Connect every stage of an A2A build — Agent Card design, auth, registry-based discovery, and production deploy — under one coherent design philosophy
✦See measured numbers for how ETag revalidation and in-process caching differ, so you know which one to reach for first
✦Understand what v1.0 renamed and restructured, and the order to migrate an existing implementation without breaking it
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
In Antigravity, create .antigravity/rules.md at the project root so A2A compliance rules are reflected in AI code suggestions:
# A2A Agent Development Rules- Every agent must serve /.well-known/agent-card.json- Task IDs must be UUID v4- Error responses must follow the A2A JSONRPC error format- Streaming must be implemented with SSE (Server-Sent Events)- Authentication must use Bearer Token (JWT)
Designing and Serving the Agent Card
We'll use a "Code Review Agent" as our running example throughout this guide.
Resilient error handling is what separates a prototype from a production system.
A2A Standard Error Codes
The A2A protocol standardizes the following JSONRPC error codes:
-32700: Parse error (JSON parse failure)
-32600: Invalid Request (missing required fields)
-32601: Method not found (unsupported method)
-32602: Invalid params (malformed parameters)
-32603: Internal error (server-side error)
-32001: Task not found
-32002: Task not cancelable
-32003: Push notification not supported
-32004: Unsupported operation
-32005: Incompatible content types
Branching on the numeric code alone gets awkward in v1.0. Per the v1.0 change notes, error payloads now follow google.rpc.Status, and the data array (details in the HTTP+JSON binding) must carry a google.rpc.ErrorInfo object. Its reason holds an UPPER_SNAKE_CASE identifier like TASK_NOT_FOUND, and domain is "a2a-protocol.org".
In practice, read reason first and fall back to the numeric code when it isn't there. The numbers belong to the JSON-RPC binding; reason comes back in the same vocabulary across all three bindings.
Exponential Backoff Retry Client
// src/retry-client.tstype RetryOptions = { maxAttempts?: number; initialDelayMs?: number; maxDelayMs?: number; retryableErrors?: number[]; // retryable A2A error codes retryOnNetworkError?: boolean; // retry dropped connections / timeouts};export async function withRetry<T>(fn: () => Promise<T>, options: RetryOptions = {}): Promise<T> { const { maxAttempts = 3, initialDelayMs = 500, maxDelayMs = 10_000, retryableErrors = [-32603], retryOnNetworkError = true, } = options; let lastError: Error | null = null; for (let attempt = 1; attempt <= maxAttempts; attempt++) { try { return await fn(); } catch (error) { lastError = error instanceof Error ? error : new Error(String(error)); if (!isRetryable(lastError, retryableErrors, retryOnNetworkError)) throw lastError; if (attempt < maxAttempts) { // Exponential backoff + jitter const delay = Math.min( initialDelayMs * Math.pow(2, attempt - 1) + Math.random() * 100, maxDelayMs ); console.warn(`Attempt ${attempt}/${maxAttempts} failed. Retrying in ${Math.round(delay)}ms...`); await new Promise(resolve => setTimeout(resolve, delay)); } } } throw lastError ?? new Error("Max retry attempts exceeded");}/** Only use the A2A code when we can actually read one. Never wave an unknown error through. */function isRetryable( err: Error, retryableErrors: number[], retryOnNetworkError: boolean): boolean { const match = err.message.match(/A2A Error (-?\d+):/); if (match) return retryableErrors.includes(Number(match[1])); // No code means this wasn't an A2A response at all: a dropped socket, an abort, or our own bug return retryOnNetworkError && isTransportError(err);}function isTransportError(err: Error): boolean { const name = err.name; return name === "TypeError" || name === "AbortError" || /fetch failed|ECONN|ETIMEDOUT/i.test(err.message);}// Usage — mint the task ID once, outside the retried callconst taskId = crypto.randomUUID();const result = await withRetry( () => client.sendTask({ taskId, message: "Please review this code" }), { maxAttempts: 3, retryableErrors: [-32603, -32000] });
Both of those corrections came from stepping on the problem myself.
The first was where taskId lives. I originally had crypto.randomUUID() inside the closure, which means every retry mints a different task ID. From the receiving agent's point of view, one request becomes three unrelated tasks piling up. Passing an ID for idempotency and then regenerating it each attempt defeats the entire point. As the pitfalls section notes below, reusing the task ID on a retry is the correct behavior.
The second was the fallback when no code could be parsed. The old version defaulted errorCode to -32603, so anything that wasn't even an A2A response — a TypeError on my side, a malformed request I'd built myself — got classified as a server internal error and retried. Retry only when the other side told you it was temporary. Blaming my own broken code on the remote agent and throwing it three more times just delayed the moment I found the bug.
Deploying to Cloudflare Workers
Antigravity pairs exceptionally well with Cloudflare Workers for edge deployment.
wrangler secret put GOOGLE_API_KEYwrangler secret put JWT_SECRETwrangler secret put AGENT_TOKENwrangler deploy# Verify the Agent Card is livecurl https://code-review-agent.your-subdomain.workers.dev/.well-known/agent-card.json | jq .
Real-World Use Case — Code Review + Documentation Pipeline
Here's a practical orchestration of three A2A agents working together:
// src/orchestrator.tsimport { A2AClient } from "./a2a-client";import { withRetry } from "./retry-client";export async function runCodePipeline(sourceCode: string) { const reviewClient = new A2AClient(process.env.REVIEW_AGENT_URL!, process.env.REVIEW_AGENT_TOKEN!); const docClient = new A2AClient(process.env.DOC_AGENT_URL!, process.env.DOC_AGENT_TOKEN!); // Verify both agents are reachable before starting const [reviewCard, docCard] = await Promise.all([ reviewClient.getAgentCard(), docClient.getAgentCard(), ]); console.log(`✓ ${reviewCard.name} (${reviewCard.version}) — online`); console.log(`✓ ${docCard.name} (${docCard.version}) — online`); // Run independent tasks in parallel to maximize throughput const [reviewResult, docResult] = await Promise.all([ withRetry(() => reviewClient.sendTask({ taskId: crypto.randomUUID(), message: `Review this code:\n\`\`\`typescript\n${sourceCode}\n\`\`\``, skillId: "review-code", })), withRetry(() => docClient.sendTask({ taskId: crypto.randomUUID(), message: `Generate JSDoc comments and a README section for:\n\`\`\`typescript\n${sourceCode}\n\`\`\``, })), ]); return { review: reviewResult.artifacts?.[0]?.parts?.[0]?.text ?? "", documentation: docResult.artifacts?.[0]?.parts?.[0]?.text ?? "", };}
Running the two tasks with Promise.all cuts wall-clock time by up to 50% compared to sequential execution. In Antigravity, add this to .antigravity/rules.md so the IDE surfaces this pattern automatically:
- Independent agent tasks must be executed with Promise.all, never sequentially
Advanced Patterns: Task Routing and Agent Discovery
As your A2A system grows, hardcoding agent URLs in your orchestrator becomes a maintenance problem. Here's a lightweight agent registry pattern that solves dynamic discovery.
Building an Agent Registry
// src/agent-registry.tstype AgentRegistration = { url: string; token: string; card?: AgentCard; lastSeen: Date; healthy: boolean;};export class AgentRegistry { private agents = new Map<string, AgentRegistration>(); register(name: string, url: string, token: string) { this.agents.set(name, { url, token, lastSeen: new Date(), healthy: true }); } async resolve(name: string): Promise<{ url: string; token: string; card: AgentCard }> { const registration = this.agents.get(name); if (!registration) throw new Error(`Agent not registered: ${name}`); if (!registration.healthy) throw new Error(`Agent unhealthy: ${name}`); // Fetch and cache the Agent Card on first use if (!registration.card) { const res = await fetch(`${registration.url}/.well-known/agent-card.json`); registration.card = await res.json(); } return { url: registration.url, token: registration.token, card: registration.card! }; } /** Find agents that support a given skill ID */ async findBySkill(skillId: string): Promise<string[]> { const matches: string[] = []; for (const [name, reg] of this.agents.entries()) { if (!reg.card) { try { const res = await fetch(`${reg.url}/.well-known/agent-card.json`); reg.card = await res.json(); } catch { reg.healthy = false; continue; } } const hasSkill = reg.card?.skills?.some(s => s.id === skillId); if (hasSkill) matches.push(name); } return matches; } /** Periodic health check — mark unresponsive agents as unhealthy */ async healthCheck() { const checks = Array.from(this.agents.entries()).map(async ([name, reg]) => { try { const controller = new AbortController(); const timeout = setTimeout(() => controller.abort(), 5000); const res = await fetch(`${reg.url}/.well-known/agent-card.json`, { signal: controller.signal, }); clearTimeout(timeout); reg.healthy = res.ok; reg.lastSeen = new Date(); } catch { reg.healthy = false; } return { name, healthy: reg.healthy }; }); return Promise.all(checks); }}// Usage in orchestratorconst registry = new AgentRegistry();registry.register("code-review", process.env.REVIEW_AGENT_URL!, process.env.REVIEW_AGENT_TOKEN!);registry.register("documentation", process.env.DOC_AGENT_URL!, process.env.DOC_AGENT_TOKEN!);registry.register("test-generator", process.env.TEST_AGENT_URL!, process.env.TEST_AGENT_TOKEN!);// Dynamically route by skillconst reviewAgents = await registry.findBySkill("review-code");console.log(`Found ${reviewAgents.length} agent(s) with review-code skill`);
Skill-Based Load Balancing
When multiple agents share the same skill (e.g., three instances of the code review agent for redundancy), a round-robin or least-loaded routing strategy prevents hotspots:
// src/load-balancer.tsexport class RoundRobinBalancer { private counters = new Map<string, number>(); next(candidates: string[]): string { if (candidates.length === 0) throw new Error("No candidates available"); const key = candidates.join("|"); const current = this.counters.get(key) ?? 0; const selected = candidates[current % candidates.length]; this.counters.set(key, current + 1); return selected; }}// In the orchestrator: pick the least-busy agent for a skillconst balancer = new RoundRobinBalancer();async function delegateToSkill(skillId: string, message: string) { const candidates = await registry.findBySkill(skillId); if (candidates.length === 0) throw new Error(`No agents available for skill: ${skillId}`); const agentName = balancer.next(candidates); const { url, token } = await registry.resolve(agentName); const client = new A2AClient(url, token); return client.sendTask({ taskId: crypto.randomUUID(), message, skillId });}
Performance Benchmarks and Optimization Tips
Understanding performance characteristics helps you make the right architectural tradeoffs.
Latency Breakdown
A typical A2A call breaks into three layers:
Network round trip — decided by distance and route, not by your code
Agent Card lookup — once per agent; what you do on the second lookup is the real design decision
Task processing — governed by the model and the prompt
For that third layer I'm deliberately not quoting seconds. The number moves by an order of magnitude depending on which model you pick, how long your prompt is, and how busy the provider is that afternoon. Carrying home a figure someone else measured on someone else's setup does not help you decide anything. Measure only the layer you can reproduce yourself. I try to hold that line both when I write numbers down and when I read my own dashboards.
So I measured the layer that has no model in it. I stood up a bare Node 22 HTTP server that serves nothing but an Agent Card, and ran 300 iterations of each access pattern against it on the same host. Network distance is deliberately near zero here, which isolates server work and serialization.
How the card is fetched
Mean
p50
p95
Body transferred
200 OK (full fetch every time)
1.94 ms
1.92 ms
2.65 ms
910 bytes
304 Not Modified (conditional, If-None-Match)
1.77 ms
1.78 ms
2.07 ms
0 bytes
In-process cache hit
0.000 ms
0.000 ms
0.000 ms
—
Building that table cost me an assumption I'd been carrying. A conditional request does not remove the round trip. Even on the same host it still took 1.77 ms, only 0.17 ms less than the full fetch — and on a real route the entire RTT sits on top of both numbers equally. From a latency point of view, a 200 and a 304 cost you nearly the same thing.
What ETag buys you is bytes and origin work, not waiting. If you want the waiting gone, you need to not make the request at all. Conditional requests reduce transfer; an in-process cache reduces round trips. They pull on different ropes, which is why one alone was never enough for me.
Optimization Strategies
1. Parallelize independent tasks aggressively. As shown in the orchestrator example, Promise.all is the single most impactful optimization. If your pipeline has any step that fans out to multiple agents with no dependencies between them, they should run in parallel.
2. Layer a TTL and an ETag together when caching Agent Cards. The common implementation is a Map with a TTL and nothing else, which means every expiry falls straight back to a full fetch. The spec's Caching Considerations ask servers to send Cache-Control and ETag, and ask clients to revalidate expired cards with a conditional request rather than re-downloading them. With both layers in place, an unchanged card costs zero bytes of body, and a changed one arrives in the same round trip you were already paying for.
type CardEntry = { card: AgentCard; etag: string | null; expiresAt: number };const cardCache = new Map<string, CardEntry>();/** Honors Cache-Control max-age; revalidates expired entries with If-None-Match */export async function getCachedAgentCard( url: string, fallbackTtlMs = 3_600_000): Promise<AgentCard> { const endpoint = `${url.replace(/\/$/, "")}/.well-known/agent-card.json`; const cached = cardCache.get(endpoint); // Still fresh: never touch the network if (cached && Date.now() < cached.expiresAt) return cached.card; const headers: Record<string, string> = {}; if (cached?.etag) headers["If-None-Match"] = cached.etag; const res = await fetch(endpoint, { headers }); // 304 carries no body — extend the copy we already hold if (res.status === 304 && cached) { cached.expiresAt = Date.now() + maxAgeMs(res, fallbackTtlMs); return cached.card; } if (!res.ok) { // A failed revalidation shouldn't take discovery down with it if (cached) return cached.card; throw new Error(`Agent card fetch failed: ${res.status}`); } const card = (await res.json()) as AgentCard; cardCache.set(endpoint, { card, etag: res.headers.get("etag"), expiresAt: Date.now() + maxAgeMs(res, fallbackTtlMs), }); return card;}function maxAgeMs(res: Response, fallbackMs: number): number { const cc = res.headers.get("cache-control") ?? ""; const m = cc.match(/max-age=(\d+)/); return m ? Number(m[1]) * 1000 : fallbackMs;}
Returning the stale card when revalidation fails is intentional. Fetching a card is discovery, not execution — throwing here makes agents that are perfectly healthy look unreachable.
3. Use streaming for tasks over 500ms. Any task that takes more than half a second benefits from streaming — it keeps your UI responsive and lets users see progress instead of staring at a spinner.
4. Tune retry behavior by the kind of failure, not one global setting. Rate limiting deserves an aggressive backoff (5–30 seconds). A server-side internal error (-32603) can retry quickly (500ms–2s). A dropped connection is worth 2–3 attempts before you surface anything to the user.
5. Keep Agent Cards small and focused — but not for the reason you'd guess. This is the other place measuring changed my mind. I grew a card with identically shaped skills and compared JSON size, gzipped size, and JSON.parse time.
Skills
JSON bytes
Gzipped
JSON.parse
3
1,234
412
0.0070 ms
5
1,768
427
0.0088 ms
10
3,103
461
0.0151 ms
20
5,793
529
0.0278 ms
Going from three skills to twenty moved the gzipped card from 412 bytes to 529. Card structure is highly repetitive, so compression absorbs almost all of it, and parsing never reaches a thirtieth of a millisecond. The reason to keep a card small is to reduce the chooser's hesitation, not the payload. Several agents with three to five skills each are easier to route to — and far easier to isolate when something breaks — than one agent carrying twenty. Argue it on payload size and you'll lose sight of the real reason.
What v1.0 Renamed — Folding an Existing Implementation Forward
The code above uses tasks/send and tasks/sendSubscribe as JSON-RPC method names. I've left them because they're the clearest shape to learn from, but the v1.0 change notes show the operation names have been reorganized. If you already have agents in the field, the mapping below makes the migration easier to size.
Through v0.3.0
v1.0
What it means for your code
message/send (this guide's tasks/send)
SendMessage
Task-vs-Message return semantics now precisely specified
message/stream (this guide's tasks/sendSubscribe)
SendStreamingMessage
kind and final removed from stream events
tasks/get
GetTask
Servers MUST return only tasks visible to the caller
tasks/cancel
CancelTask
Cancellable state transitions now spelled out
tasks/resubscribe
SubscribeToTask
Reconnection and subscription lifecycle formalized
(none)
ListTasks
New, with filtering and cursor-based pagination
The structural changes cost more than the renames did. Stream events dropped the kind discriminator, so TaskStatusUpdateEvent and TaskArtifactUpdateEvent are now told apart by JSON member name. Enum values moved from kebab-case to SCREAMING_SNAKE_CASE. On the Agent Card, protocolVersion moved off the card and onto each AgentInterface, while preferredTransport and additionalInterfaces were consolidated into supportedInterfaces[]. The HTTP+JSON binding also dropped its /v1 URL prefix.
Sequencing matters here: make the receiving side accept both shapes first, then switch what you send. Get it receivable before you change how you send. I apply that order to anything with two ends, not just protocol upgrades — reverse it and you spend the gap emitting messages nobody on the other side can read.
Common Pitfalls and How to Avoid Them
Building A2A systems for the first time surfaces a few recurring mistakes. Here are the ones worth knowing about upfront.
Pitfall 1: Forgetting to set aud in JWT tokens. If your JWT doesn't specify the intended audience (aud), any compromised token can be replayed against any agent in your system. Always include aud: targetAgentId and verify it on the receiving end.
Pitfall 2: Not handling partial streaming artifacts. When reconnecting to a streaming task after a dropped connection, the client may receive overlapping chunks. Track the artifact.index field and skip chunks with an index lower than the last one you successfully processed.
Pitfall 3: Using the same task ID for different tasks. Task IDs must be unique per logical unit of work. Reusing a task ID for a retry is correct (that's exactly why the retry client above hoists the ID out of the closure), but using it for a semantically different request causes unpredictable behavior in agents that cache task state.
Pitfall 4: Serving Agent Cards without caching headers. Without Cache-Control: public, max-age=3600, every agent discovery call hits your Worker's compute budget. This adds up quickly when you have many orchestrators polling for available agents. Send an ETag alongside it — the spec suggests deriving one from the card's version field or a content hash, which is what lets clients settle for a conditional request.
Pitfall 5: Running health checks too frequently. Aggressive health checking (every 5 seconds across 20 agents) creates unnecessary load. A 30-second interval with a 5-second timeout is a reasonable starting point for most production systems.
What I Learned Running This — Where Standard Protocols Pay Off
Run an indie app long enough and you will, at some point, build your own way of handing background jobs to a pool of workers. I did. And what hurt was never the features — it was maintaining the bespoke protocol I'd invented. Sender and receiver schemas drift apart quietly, and fixing one side breaks the other. Nobody is at fault, and the thing you repaired catches fire again the following week.
That unglamorous maintenance reality is exactly where a standard like A2A earns its keep, by putting capability declaration (the Agent Card), task hand-off, and error semantics outside your codebase. Having the spec live somewhere else also means you argue with a document instead of a colleague, which I may value more than any of its technical properties.
Three judgment calls became clear only after running this in practice.
First, Agent Card caching gets framed as a speed optimization, but what it really buys you is failure isolation. The measurements above show how little latency it saves in some configurations. I still put it in every time, because a cached card lets me split a failed task in the logs immediately into "discovery failed" versus "processing failed." What I'd installed as a speed mechanism turned out to be a diagnostic one — I had the order backwards, and only noticed at the end of a long night chasing a cause.
Second, set your health-check interval from the average task duration, not the number of agents. Running 5-second checks against a system full of long-running tasks will flag perfectly healthy agents as unhealthy on a transient timeout. The instinct to tighten the interval as agent count grows usually pushes in the wrong direction.
Third, don't build the perfect discovery mechanism up front. While you have three or four agents, a registry as lightweight as the one in this guide is plenty. Premature abstraction only ties your hands later. Make the registry smarter once manual routing genuinely starts to annoy you — that moment arrives on its own.
Summary
In this guide, we built a production-grade A2A-compliant agent system with Antigravity — from Agent Card design all the way to Cloudflare Workers deployment.
Key takeaways:
Agent Cards are served at /.well-known/agent-card.json and declare capabilities, auth, and skills
Task delegation has three patterns: synchronous, streaming, and push notifications — choose based on expected latency
JWTs must include iss, aud, and scopes to identify and authorize agents precisely
Error handling should follow A2A standard error codes with exponential backoff for retryable failures
Cloudflare Workers is the ideal deployment target for A2A agents — edge distribution with minimal latency
Parallel task execution via Promise.all is the single biggest throughput win in orchestrator agents
Conditional requests and in-process caching solve different problems — transfer size versus round trips — so layer both
v1.0 renamed the operations and restructured the card, so make the receiving side tolerant before you change what you send
Where the spec itself is still moving, reach for the primary source before you reach for the code. That habit is what I leaned on hardest while revising this guide. Thank you for reading this far.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.