SRE for Antigravity Agents — Taming Probabilistic Systems with SLOs and Error Budgets
AI agents are probabilistic by nature, so running them in production without SRE thinking is risky. This guide applies SLIs, SLOs, and error budgets to Antigravity agents with working code, and shows why the textbook 14.4x burn-rate alert never fires on a loose SLO.
I had started handing the first draft of app-store review replies to an agent. For weeks the dashboard stayed green, yet most mornings I opened the drafts and found something I had to rewrite.
HTTP 200s were coming back. The error rate was close to zero. And the one number I cared about — how many drafts were actually usable — appeared nowhere on the screen. I still remember the heavy feeling in my stomach.
What I'd like to say first: a probabilistic agent can be run in production with the classic SRE toolkit — SLIs, SLOs, and error budgets. But if you define the SLI the way you would for a web service, all you get is a green screen. This guide walks through how I define them, the instrumentation code, and a trap I noticed while checking the math: an alert that is configured correctly and can never fire.
The classic SLIs in the Google SRE Book — HTTP success rate, latency under a threshold — don't translate cleanly. If an agent returns a 200 for "refactor this module," the status code says nothing about whether the refactor makes sense. You won't know until you read the diff.
My first instinct was is_success = response.status == 200. With that definition, an agent that returns a nonsensical diff still counts as a success. The metrics look perfect while the work for the person using it doesn't shrink — that gap is the first trap.
Here is a second one that is easy to fall into. Define "completed" as "the agent called the submit_result tool," and the agent can satisfy the SLI by calling it with empty arguments. The agent isn't being devious. When you measure what's convenient, optimization drifts toward the artifact.
Three lines I now draw:
Separate deterministic success from probabilistic quality — you need both, and neither alone is enough
Measure what the user experiences (was the task done?), not internal API status codes
Assume sampled evaluation, because running every output through an LLM judge gets expensive fast
Changing an SLI definition later means rebuilding everything stacked on top of it: alerts, runbooks, agreements. Spending a week on the definition before writing any code is not wasted time.
Define SLIs in Three Layers
I use three layers, each measuring a different facet of reliability.
Tool-call success, schema match, type check, tests passing
Plain programs (deterministic)
Every request
3. Quality
Requirement fit, convention fit, hallucination
LLM-as-a-Judge plus human review
Sampled (a few % to 10%)
Layer 1 alone isn't enough, but if it's broken nothing above it matters, so cover it first. Layer 2 needs no LLM, so instrument every request. Layer 3 is expensive and belongs in a separate job.
If you skip this split and try to measure everything with a judge, costs balloon and evaluation noise destabilizes your SLO. I got the order wrong once, and I'd rather you didn't.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦Turn 'AI agents are too unpredictable for production' from a blocker into a measurable reliability target you can commit to
✦Copy an instrumentation stack that combines deterministic checks with LLM-as-a-Judge, and drop it into your own project
✦Work backwards from your SLO to the burn rates that can actually fire, so you never ship an alert that stays silent forever
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
The typecheck function can simply write the files to a temp directory and run tsc --noEmit. I once nearly settled for "does the content contain the word export," which lets near-empty output through. Start with a completion check that's slightly too strict; it's easier to loosen later.
Write one verify function per task type. "Done" means something different for code generation, doc generation, and research reports, so keep them pluggable.
Instrumenting Layer 3 (LLM-as-a-Judge)
Quality SLIs run as a separate job. For cost, latency, and reliability alike, keeping them off the request path is safer.
// quality-judge.ts// Purpose: score sampled outputs with an LLM judge// Key idea: if the judge fails, production must notimport { GoogleGenAI } from "@google/genai";const judge = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });const JUDGE_PROMPT = `You evaluate the output quality of an AI agent.Score the item below on three axes from 1 to 5 and answer in JSON.- correctness: does it satisfy the task requirements?- safety: any security problems (hard-coded secrets, SQL injection, etc.)?- style: does it follow the project's coding conventions?Task: {task}User input: {input}Agent output: {output}Return: {"correctness": 1-5, "safety": 1-5, "style": 1-5, "reason": "..."}`;interface QualityScore { correctness: number; safety: number; style: number; reason: string;}export async function evaluateQuality( task: string, input: string, output: string,): Promise<QualityScore | null> { // Use function replacers so "$&" and similar in the input cannot corrupt the prompt const prompt = JUDGE_PROMPT .replace("{task}", () => task) .replace("{input}", () => input.slice(0, 4000)) .replace("{output}", () => output.slice(0, 8000)); try { const response = await judge.models.generateContent({ model: "gemini-2.5-flash", contents: prompt, config: { responseMimeType: "application/json", temperature: 0.1, // keep scoring stable }, }); const text = response.text; if (!text) { console.warn("judge returned empty response"); return null; } const parsed = JSON.parse(text) as QualityScore; const scores = [parsed.correctness, parsed.safety, parsed.style]; if (!scores.every((v) => typeof v === "number" && v >= 1 && v <= 5)) { console.warn("judge returned invalid scores", parsed); return null; } return parsed; } catch (err) { // A judge failure is not an agent failure. Log it and move on. console.error("judge failed:", err); return null; }}
Plain random sampling isn't enough either, because rare task types never accumulate evaluations. Guarantee a floor per type.
// sampler.ts// Purpose: stratified sampling with a minimum per task typeconst evaluatedCount = new Map<string, number>();const MIN_PER_TYPE_PER_DAY = 5;const BASE_RATE = 0.05;export function shouldSample(taskType: string): boolean { const done = evaluatedCount.get(taskType) ?? 0; const pick = done < MIN_PER_TYPE_PER_DAY || Math.random() < BASE_RATE; if (pick) evaluatedCount.set(taskType, done + 1); return pick;}// Call at the start of the daily jobexport function resetDailyCounters(): void { evaluatedCount.clear();}
People ask whether the judge's scores can be trusted. My answer is "not fully." The judge is a coarse sieve for large volumes; final quality calls get corrected by a human. Even one hour a week of reading what the judge scored low, and measuring how often your scores agree with it, goes a long way. Track that agreement rate as a meta-SLI.
Choosing SLOs Without Emotion
Once the SLIs exist, set the SLOs. "Just say 99.9%" is the trap. Work backwards from the failure the business can tolerate.
Here is an example for a code-generation agent. These numbers illustrate the reasoning; they are not values to adopt as they are.
Layer
Example SLO
Why
1
99.5% monthly request success, P95 under 5 s
Borrow the standard from conventional services
2
92% monthly task completion
Some requests are ambiguous enough that a human would get them wrong too
3
Sampled correctness mean ≥ 4.0, safety mean ≥ 4.5
Safety errors cost more, so the bar is higher
I don't put 99.9% on layer 2 because a target you can't reach only wears out whoever runs it. Start slightly looser than what you measure over one or two weeks of real traffic, then draw the line at "a number we can keep."
An Error-Budget Gate
An SLO implies its flip side, the error budget. A 92% monthly SLO leaves 8% of failures you can afford.
The budget's value is that decisions change with what's left: ship freely while it's plentiful, hold risky changes when it's nearly gone. This code automates that switch.
// error-budget-gate.ts// Purpose: decide whether to deploy or experiment, based on remaining budget// Key idea: leave the call to the numbers, not to a debateexport interface ErrorBudgetStatus { consumedPercentage: number; // 0 or more; above 100 means overspent remainingPercentage: number; // never below 0 recommendation: "proceed" | "proceed_with_caution" | "freeze"; reason: string;}const QUERY = ` sum(rate(agent_task_completions_total{outcome="completed"}[28d])) / sum(rate(agent_task_completions_total[28d]))`;export async function checkErrorBudget( promUrl: string, sloTarget: number, // e.g. 0.92): Promise<ErrorBudgetStatus> { const res = await fetch( `${promUrl}/api/v1/query?query=${encodeURIComponent(QUERY)}`, { signal: AbortSignal.timeout(5000) }, ); if (!res.ok) { // How to treat a Prometheus outage is a policy choice; some teams should freeze return { consumedPercentage: 0, remainingPercentage: 100, recommendation: "proceed_with_caution", reason: `prometheus query failed: ${res.status}`, }; } const data = (await res.json()) as { data: { result: Array<{ value: [number, string] }> }; }; // An empty result (just instrumented, no traffic) or NaN must not read as "0% used" const raw = data.data.result[0]?.value[1]; const completionRate = raw === undefined ? NaN : parseFloat(raw); if (Number.isNaN(completionRate)) { return { consumedPercentage: 0, remainingPercentage: 100, recommendation: "proceed_with_caution", reason: "no data in window — cannot evaluate budget", }; } const allowedFailure = 1 - sloTarget; // e.g. 0.08 const actualFailure = 1 - completionRate; const consumed = (actualFailure / allowedFailure) * 100; const remaining = Math.max(0, 100 - consumed); if (consumed < 50) { return { consumedPercentage: consumed, remainingPercentage: remaining, recommendation: "proceed", reason: "budget has plenty of headroom" }; } if (consumed < 90) { return { consumedPercentage: consumed, remainingPercentage: remaining, recommendation: "proceed_with_caution", reason: "budget is being consumed — risky changes require review" }; } return { consumedPercentage: consumed, remainingPercentage: remaining, recommendation: "freeze", reason: "error budget nearly exhausted — freeze non-critical deploys" };}
Keep the caller in its own file. This assumes an ESM runner such as tsx.
Call it from GitHub Actions and fail the job on freeze, and the mood of operations changes. Nobody has to wake up at night and plead for a freeze; the pipeline stops first.
"Prometheus is down, so proceed with caution" is a real judgment call. I judged that agents halting on a short outage costs more than a bad deploy slipping through that window. Your answer may differ. Write the reasoning down as an ADR, so the person you'll be in six months can question the line.
Burn-Rate Alerts That Can Actually Fire
Alert design is where agent SRE differs most. An alert that fires the instant a failure happens wears the on-call out. The Google SRE Workbook's multi-window, multi-burn-rate approach applies here.
Burn rate is how many times faster than the sustainable pace you're spending the budget. A burn rate of 1 uses the budget up exactly at the end of the period.
Here is the trap I noticed while checking the math. The Workbook's standard example — burn rate 14.4 over one hour, 6 over six hours — assumes an SLO around 99.9%. Carry it over to a 92% agent SLO and this happens:
Setting
Burn rate
Window
Budget consumed (30 days)
Failure ratio needed (SLO 92%)
Textbook fast
14.4
1 hour
2%
115.2% (unreachable)
Textbook slow
6
6 hours
5%
48%
This guide: page
6
1 hour
about 0.8%
48%
This guide: ticket
3
6 hours
2.5%
24%
This guide: slow
1
3 days
10%
8%
With an allowed failure rate of 8%, the failure ratio can never exceed 100%. So the highest burn rate that can ever fire is 1 ÷ 0.08 = 12.5, and a 14.4x alert will never page you however long you wait. Your dashboard stays green and the alert stays silent — the hardest kind of failure to notice.
You can verify this in a few lines of Node. Rebuild this table every time you change an SLO.
Run with SLO 0.92, it prints the same numbers as the table: a maximum of 12.50, page at 48.0%, ticket at 24.0%, slow at 8.0%.
Now the alert itself. Fire only when both a short and a long window exceed the threshold. The long window confirms the spend is real; the short one confirms it is still happening. First, record the per-window failure ratios.
# agent-slo-rules.yml (excerpt; define 1h, 6h, 30m, and 3d the same way)groups: - name: agent-slo-recording rules: - record: agent:task_failure_ratio:rate5m expr: | 1 - ( sum(rate(agent_task_completions_total{outcome="completed"}[5m])) / sum(rate(agent_task_completions_total[5m])) ) - record: agent:task_failure_ratio:rate1h expr: | 1 - ( sum(rate(agent_task_completions_total{outcome="completed"}[1h])) / sum(rate(agent_task_completions_total[1h])) ) - name: agent-slo-alerts rules: - alert: AgentErrorBudgetBurnPage expr: | agent:task_failure_ratio:rate1h > (6 * 0.08) and agent:task_failure_ratio:rate5m > (6 * 0.08) and sum(rate(agent_task_completions_total[1h])) > 0.01 for: 2m labels: severity: page annotations: summary: "At this pace the budget is gone in about 5 days (SLO 92%)"
The trailing > 0.01 is a traffic floor against false alarms in quiet hours. An agent handling about 36 requests an hour will see its 5-minute failure ratio swing wildly on a single failure. The lower your volume, the more this floor and your window lengths matter.
One more note. The budget math above assumes 30 days. If your SLO window is 28 days, convert the consumed-budget figures with 28 days; when the window and the burn rate drift apart, the alert stops matching the promise you made.
For the first month, I'd set the thresholds looser (page at 4 instead of 6, say) and tighten them as you watch the firing rate. If a week passes without a page, don't rush to lower the threshold. An alert that stays quiet is working.
Incident Patterns Worth a Runbook
Having SLOs doesn't tell you what to do when something breaks. Three agent-specific patterns make a good first draft of a runbook.
Pattern 1: a sudden spike in failures (10% to 50% within an hour)
Suspect the upstream LLM provider first: check its status page and your own latency metrics for anomalies, both sudden rises and suspicious drops. Then suspect the last prompt or tool-definition change. Prompts are fragile; one line can change behavior. Keep them in Git like code and write the rollback steps into the runbook.
Pattern 2: quality that decays slowly (the budget leaks over days)
The hardest one. Infra and layer 2 look fine while judge scores slide from 4.2 to 4.0 to 3.8. Usually it is data drift — inputs moving away from what the prompt was designed for. Record token-length distributions and library versions over time so you can pin down when it started.
Pattern 3: failures isolated to one task type
Overall SLIs look healthy, yet one segment suffers. If you label by task type, you can carve a segment-specific SLO out of the same metrics. Don't be stingy with labels; they pay off later.
Evaluation Datasets as First-Class Artifacts
The credibility of a quality SLI rests on its evaluation dataset. Keep input examples paired with a checklist of what a good output satisfies — not an exact expected answer. For code generation: "imports React," "handles the empty array," "follows the naming convention."
Keep the dataset in Git and tag every significant change; let judge prompts reference checklist items. You get three things: you can re-run last month's evaluation against this month's agent and compare, you can diff your way to a cause when quality drops, and non-engineers can read the checklist and ask useful questions.
Refresh it every month or so with samples of real inputs, after consent and a privacy review. Replace only the oldest third; the stable examples are what let you notice a regression. Sampled inputs may contain personal data or secrets, so route them through the same classification and redaction as your logs and keep them out of public repositories.
On a team, present the SLO as a commitment line for service quality. Agreeing in advance that "below 92% completion we pause new features and work on the agent" keeps quality-versus-revenue from turning into a tug of war. For feature developers, show the budget as a budget: "we've used 60% this month, so let's push risky prompt changes to next month."
If you run apps or sites on your own, the person you have to agree with is yourself. Even so, writing the SLO into a README, with "below this number I stop shipping features," helps. A tired me at night tends to make the lenient call.
Build two dashboards: a one-page overview, and a time-series view for digging in. One screen trying to do both is awkward for everyone.
Label Design
The labels on layer 2 pay off more the longer you run.
Label
Purpose
task_type
Carve out SLOs per task type
tenant_id
Catch regressions for one customer or team (skip if you run alone)
model
Compare models and decide on switches (use the model ID as is)
prompt_version
Hash of the prompt template, to trace the effect of changes
client_version
Isolate client-side regressions
Don't label by user_id or anything else high-cardinality; per-user investigation belongs to structured logs and traces. Metrics aggregate, logs and traces individuate. Before adding a label, imagine two or three concrete queries you'd write with it. If you can't, you don't need it yet.
When to Skip SLOs
Not every agent needs this machinery. Under about 100 requests a day, percentage-based SLOs are statistical noise; structured logs and a human eye are enough. A prototype whose purpose is still shifting is the same, since an SLO would lock in assumptions you're about to change.
If you already have a product-level SLO such as "90% of sessions end in a successful outcome" and the agent's behavior is inside it, a separate agent SLO may be unnecessary. But if the agent touches money, handles security-sensitive operations, or runs unsupervised, the cost of an SLO is far smaller than one bad week without it. Instrument from day one.
Pitfalls
These are ones I've hit, or nearly hit.
Pitfall 1: keeping the SLO from the people it affects. If only operations knows the SLO, a request like "let's just ship this speed optimization" quietly eats the budget. Publish the remaining budget somewhere anyone can see it.
Pitfall 2: not versioning the evaluation dataset. When judge scores drop, you can't tell whether the agent regressed or the dataset changed. Tag metrics with the Git commit ID.
Pitfall 3: alerting on the SLI directly. A page at 5% failures wears people out. Alert on burn rate, and make sure the threshold is reachable from your SLO.
Pitfall 4: deferring human review. The judge has tastes of its own. Without an hour of human eyes a week, a broken judge goes unnoticed. Count that hour as an SRE cost.
Pitfall 5: one SLO for many agents. One agent's regression hides inside another's surplus. Give each its own SLO and dashboard.
The First Step
If you'd like to start, please don't build everything at once. SRE comes in at the pace of your maturity.
Instrument layer 2 only — task completion rate — for one week. Layer 1 is usually covered by existing monitoring, and layer 3 costs both money and effort. When the number arrives, notice whether it is higher or lower than you expected; that reaction is where the SLO conversation starts. Then, once you've chosen an SLO, put it into burn-table.mjs and check, just once, that the alert you configured can actually fire.
That is the order in which I rebuilt mine. I still treat the quiet alerts as the ones most worth doubting.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.