◉ANTIGRAVITY LABJP
Articles/Agents & Manager
◈ Agents & Manager/2026-04-24Advanced

SRE for Antigravity Agents — Taming Probabilistic Systems with SLOs and Error Budgets

AI agents are probabilistic by nature, so running them in production without SRE thinking is risky. This guide applies SLIs, SLOs, and error budgets to Antigravity agents with working code, and shows why the textbook 14.4x burn-rate alert never fires on a loose SLO.

antigravity461sreslo3error-budgetproduction71reliability12agents144

✦ Premium Article

I had started handing the first draft of app-store review replies to an agent. For weeks the dashboard stayed green, yet most mornings I opened the drafts and found something I had to rewrite.

HTTP 200s were coming back. The error rate was close to zero. And the one number I cared about — how many drafts were actually usable — appeared nowhere on the screen. I still remember the heavy feeling in my stomach.

What I'd like to say first: a probabilistic agent can be run in production with the classic SRE toolkit — SLIs, SLOs, and error budgets. But if you define the SLI the way you would for a web service, all you get is a green screen. This guide walks through how I define them, the instrumentation code, and a trap I noticed while checking the math: an alert that is configured correctly and can never fire.

It pairs well with the AI model evaluation and testing framework and the LangFuse observability guide. Read together, they cover the whole loop from instrumentation to decisions.

Why Traditional SLIs Break for AI Agents

The classic SLIs in the Google SRE Book — HTTP success rate, latency under a threshold — don't translate cleanly. If an agent returns a 200 for "refactor this module," the status code says nothing about whether the refactor makes sense. You won't know until you read the diff.

My first instinct was is_success = response.status == 200. With that definition, an agent that returns a nonsensical diff still counts as a success. The metrics look perfect while the work for the person using it doesn't shrink — that gap is the first trap.

Here is a second one that is easy to fall into. Define "completed" as "the agent called the submit_result tool," and the agent can satisfy the SLI by calling it with empty arguments. The agent isn't being devious. When you measure what's convenient, optimization drifts toward the artifact.

Three lines I now draw:

  • Separate deterministic success from probabilistic quality — you need both, and neither alone is enough
  • Measure what the user experiences (was the task done?), not internal API status codes
  • Assume sampled evaluation, because running every output through an LLM judge gets expensive fast

Changing an SLI definition later means rebuilding everything stacked on top of it: alerts, runbooks, agreements. Spending a week on the definition before writing any code is not wasted time.

Define SLIs in Three Layers

I use three layers, each measuring a different facet of reliability.

LayerWhat it measuresHow it is judgedCoverage
1. InfrastructureRequest success rate, P95/P99 latency, timeout rateSame as ordinary web monitoringEvery request
2. Task completionTool-call success, schema match, type check, tests passingPlain programs (deterministic)Every request
3. QualityRequirement fit, convention fit, hallucinationLLM-as-a-Judge plus human reviewSampled (a few % to 10%)

Layer 1 alone isn't enough, but if it's broken nothing above it matters, so cover it first. Layer 2 needs no LLM, so instrument every request. Layer 3 is expensive and belongs in a separate job.

If you skip this split and try to measure everything with a judge, costs balloon and evaluation noise destabilizes your SLO. I got the order wrong once, and I'd rather you didn't.

✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦Turn 'AI agents are too unpredictable for production' from a blocker into a measurable reliability target you can commit to
✦Copy an instrumentation stack that combines deterministic checks with LLM-as-a-Judge, and drop it into your own project
✦Work backwards from your SLO to the burn rates that can actually fire, so you never ship an alert that stays silent forever
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

◈ Agents & Manager2026-05-11
Canary Deployment with Auto-Rollback for AI Agents — Protecting Production with Antigravity and Burn-Rate SLOs
A practical playbook for shipping new AI agent versions through canary deployment on Antigravity, with automatic rollback driven by burn-rate SLOs. Includes a lightweight setup that solo developers can sustain.
◈ Agents & Manager2026-04-28
Designing Production Incident Runbooks for Antigravity Agents: A Practical Framework from Detection to Recovery
A complete guide to designing incident runbooks for production Antigravity Agents — detection, triage, mitigation, and postmortem, with working code you can drop into your stack today.
◈ Agents & Manager2026-05-27
Record & Replay for Antigravity Agents
How to deterministically replay a failed Antigravity Agent run offline, drawn from a month of running it across four production sites. Covers boundary recording, R2 + KV storage costs, PII masking, and a working TypeScript harness.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links