◉ANTIGRAVITY LABJP
Articles/Agents & Manager
◈ Agents & Manager/2026-05-27Advanced

Record & Replay for Antigravity Agents

How to deterministically replay a failed Antigravity Agent run offline, drawn from a month of running it across four production sites. Covers boundary recording, R2 + KV storage costs, PII masking, and a working TypeScript harness.

antigravity461agents144observability19record-replayproduction71debugging15

✦ Premium Article

The agent went quiet late on a Sunday night. By Monday morning the logs were open on my screen and the last line before the silence was different every single time, with no obvious place to start reading. Running several Lab sites in parallel as a solo developer, that wish — if only I could see it again under the same conditions — comes up most weeks. What I want to write down here is the machinery that stops you standing there empty-handed: recording an agent's production run so you can replay it offline, deterministically, afterwards.

A small Monday-morning oddity that logs couldn't explain

A few weeks ago, one of my overnight automations started failing roughly one out of every three Sunday-night runs. GitHub would contain the first 60% of its output and nothing after — possibly the worst kind of half-failure, because the symptom is identical from outside but the cause is somewhere deep inside the agent's loop.

When I opened the logs, the last utterance before the agent went silent was different every time. Once it was mid-tool-call. Once it was mid-stream of a model response. The only common thread was that the silence began near a moment of waiting on an external dependency — the model API, GitHub, or Cloudflare KV.

In other words, no amount of reading after the fact let me reconstruct what the agent had seen at that instant, or in what order events had unfolded. Failures you cannot reproduce are, as a rule, failures you cannot fix. This article describes the Record & Replay design I built to escape that situation.

Why "just read the logs" isn't enough

Failures in autonomous agents differ in nature from classic stateless request/response outages, for three reasons.

First, agent state is the accumulated product of nondeterministic interactions between model output and tool output. Even with the same input prompt, the order and phrasing of the model's tool calls drift slightly, and the tool returns drift with them.

Second, agents carry long-form context. Hand the Antigravity Agent Manager more than ten steps of work and your context window quickly grows into tens of thousands of tokens. It is utterly routine for the eventual failure to be rooted in a tool return seven steps earlier.

Third, external side effects often cannot be rolled back. A pushed commit, a created Stripe customer, a written KV key. In production, "just rerun it under the same conditions" is rarely available.

Record & Replay attacks all three at once. Concretely:

  • Record: during the production run, capture every input and output that crosses an external boundary — model, tool, time — as a structured trace.
  • Replay: load the trace and rerun the agent locally in a deterministic loop. External APIs are not invoked; recorded returns are served by stubs instead.
  • Bisect: slice the trace around the failure point and binary-search to find the originating step.
✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦A working TypeScript harness that lets you replay a failed agent run locally in under three minutes.
✦Real cost numbers (about ¥48 / $0.30 per month for four parallel sites) for a Cloudflare R2 + KV trace store, with TTL design.
✦Concrete pitfalls — PII masking, model determinism, timezone drift — surfaced from a month of real four-site operation.
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

◈ Agents & Manager2026-05-12
Tracking AI Agent Decisions in Antigravity: Implementing Decision Logs and Explainability
Learn how to design decision logging systems that capture why your AI agents make specific choices. Includes working Python code examples, Before/After patterns, and a quality improvement cycle for production Antigravity agents.
◈ Agents & Manager2026-05-11
Canary Deployment with Auto-Rollback for AI Agents — Protecting Production with Antigravity and Burn-Rate SLOs
A practical playbook for shipping new AI agent versions through canary deployment on Antigravity, with automatic rollback driven by burn-rate SLOs. Includes a lightweight setup that solo developers can sustain.
◈ Agents & Manager2026-05-10
Giving Your Antigravity AI Agents a 'Time Budget' — A Production Scheduling Design That Unifies Timeouts, Priorities, and Deadlines
Your AI agent's response time creeps up, users drop off, and unexpected costs pile on. Here is how I give Antigravity agents a single 'Time Budget' object that unifies timeouts, priority, and deadlines, plus the traps I hit after half a year in production.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links