Where to Start Reading an Unattended Agent's Changes — A Digest for Re-Entry
How do you review the pile of changes an unattended agent left overnight? Not the full diff, not the chat log — a re-entry digest built from rule-based risk classification, per-agent review markers, and a 30-second verification pass.
An agent I ran unattended had rewritten nearly 40 files overnight. I opened it first thing in the morning, tried to read every diff top to bottom, and gave up almost immediately. The chat log was long too, and just tracing where each decision was made dissolved the time.
When you run several apps and sites in parallel in indie development, this "next-morning review" of unattended runs becomes a daily task. The problem is not the volume of changes — it is the cost of a human re-entering that context. Here I design a re-entry digest that groups changes by risk class, so you never have to read the full diff top to bottom.
The Enemy Is Re-Entry Cost, Not Volume
Reviewing an unattended run is painful not because the diff is large, but because you cannot tell where it is safe to start. Of 40 files, only a few truly deserve a close look; the rest is often confirmation-free noise like formatting or renames.
But a diff presents the critical line and a whitespace tweak with the same appearance. A human reads them top to bottom and cannot hold focus to the end. What you need is a layer that reorders changes into "the order a human should review them."
There is a second trap I fell into for a long time: scrolling back through the chat log whenever something felt unclear. An agent's reasoning log is useful for finding out why a decision was made, but it is the wrong place to confirm what actually happened. Mixing the record of reasons with the record of results is what makes re-entry expensive. What you should read is the result — the repository state and the side effects that left the machine.
Group by Risk Class
I group changes into four classes. The reading order runs from the hardest to undo.
Class
What it includes
Reading order
Depth of check
Irreversible / external
push, deploy, billing, data deletion, writes to external APIs
First
The real thing, line by line
Contract
public API, schema, config values, dependency versions
Next
Reason for change and backward compatibility
Internal
function bodies, tests, refactors
After that
Confirm tests pass
Noise
formatting, renames, comments, import reordering
Last (skippable)
Count only
This reordering alone changes how the review feels. It lets you spend your first few minutes — when focus is highest — on the changes that are hardest to take back. A single line touching AdMob config or billing belongs at the head of "irreversible / external."
What decides whether this table is actually usable is how you handle the edge cases. A CI workflow definition looks like internal code but is really the entrance to an irreversible deploy. A lockfile carries a huge line count yet belongs to the contract layer, because it decides what your dependencies actually are. A migration changes class entirely depending on whether it has already been applied.
I keep only two tie-breaking rules. When in doubt, promote one class up. And any change that includes a deletion gets promoted one class regardless of type. Erring toward the side that needs restoring keeps the damage small when the classifier is wrong.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦A rule table for classifying changes by risk class, plus a working script that does the sorting for you
✦A per-agent 'last review point' that survives parallel runs, rebases, and rewritten history
✦Machine-generated facts, agent-written reasons, and a one-command verification pass for irreversible operations
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
If you judge the class by hand every morning, that judgment itself becomes re-entry cost. Let path and filename rules do the sorting, and let the human only look at the result. The script below reads the changes since your last review point and sorts them into the four classes.
// scripts/classify-changes.mjs// usage: node scripts/classify-changes.mjs reviewed/antigravityimport { execSync } from "node:child_process";const base = process.argv[2] ?? "HEAD~1";// Evaluated top to bottom; the first matching rule winsconst RULES = [ { cls: "irreversible", re: /^(\.github\/workflows\/|wrangler\.toml$|.*\/migrations\/|infra\/|scripts\/deploy)/ }, { cls: "irreversible", re: /(stripe|billing|webhook|checkout)/i }, { cls: "contract", re: /^(src\/config\/|.*\/api\/|.*\.d\.ts$|package(-lock)?\.json$|.*schema.*\.(sql|ts|json)$)/ }, { cls: "internal", re: /^(src\/|app\/|lib\/|tests?\/)/ }, { cls: "noise", re: /(\.md$|\.snap$|\.lock$|^public\/)/ },];const ORDER = ["irreversible", "contract", "internal", "noise"];const UP = { noise: "internal", internal: "contract", contract: "irreversible", irreversible: "irreversible" };const raw = execSync(`git diff --name-status ${base}..HEAD`, { encoding: "utf8" }).trim();const buckets = { irreversible: [], contract: [], internal: [], noise: [] };for (const line of raw ? raw.split("\n") : []) { const [status, ...paths] = line.split("\t"); const path = paths.at(-1); let cls = RULES.find((r) => r.re.test(path))?.cls ?? "contract"; // unknown paths land mid-table if (status.startsWith("D") || status.startsWith("R")) cls = UP[cls]; // promote deletes and renames buckets[cls].push(`${status}\t${path}`);}for (const cls of ORDER) { const files = buckets[cls]; if (cls === "noise") { console.log(`\n## noise: ${files.length} files (count only)`); continue; } console.log(`\n## ${cls}: ${files.length} files`); files.forEach((f) => console.log(" " + f));}
Rules are evaluated in order with first-match-wins so that the classification stays explainable after the fact. When a classification is wrong, what you fix is one line of a rule, not that morning's judgment. Sending unknown paths to contract follows the same reasoning: putting the unfamiliar in a safe middle class means nothing gets silently skipped.
The noise class prints a count and nothing else. If you print the contents, you will read them anyway. Keeping what you intend to skip out of your field of view is the whole job of this layer.
Keep a Review Marker Per Agent
Reviewing the entire set of changes every morning is wasted effort. What you should look at is only the diff "after the point you last reviewed." So leave a marker at the point where you finished reviewing.
# When you finish a review, leave a marker at that pointgit tag -f reviewed/antigravity HEADgit push -f origin reviewed/antigravity# Next morning, look only at what came after the last review pointgit diff reviewed/antigravity..HEAD --statgit log reviewed/antigravity..HEAD --oneline
By moving the tag forward as your last-review point, no matter how many times the unattended run fires, a human always faces only "the diff since last time." The smaller you keep that base, the lighter the re-grouping by risk class becomes.
This broke the moment I started running agents in parallel. With a single marker, reviewing one agent and advancing the tag pushes the other agent's unreviewed work behind the marker, where it disappears from view. That is exactly how I once missed a small config change.
The fix is simply to keep one marker per agent. Name them reviewed/nightly-refactor, reviewed/link-audit, and scope each review to that agent's own commits.
AGENT="nightly-refactor"BASE=$(git rev-parse -q --verify "reviewed/${AGENT}" || git merge-base origin/main HEAD)# Restrict the review to commits from that agentgit log --committer="agent/${AGENT}" --oneline "${BASE}..HEAD"node scripts/classify-changes.mjs "$BASE"
Falling back to merge-base when git rev-parse -q --verify fails is insurance for the case where a rebase or history rewrite left the marker pointing at a commit that no longer exists. The one genuinely scary failure mode here is a broken marker that quietly reports "no changes." Recovery days produce a large diff, which is far better than missing one.
Have the Agent Write the Digest by "Perspective"
The change summary itself is fine to have the agent write. But bullet-listing "what it changed" only paraphrases the diff. What you should make it write is the human's review perspective.
At the end of the run, always have it output three things. First, a list of operations in the irreversible / external class (state "none" explicitly if there are none). Second, changes that could break backward compatibility at the contract layer, with reasons. Third, the files and lines a human should look at first.
Rather than pinning the output format, pin the perspective in the agent's instructions. I keep this fragment at the end of every unattended task.
End the run with a "re-entry digest".- List every operation in the irreversible / external class (push, deploy, billing, deletion, writes to external APIs). If there were none, say "none" explicitly. Do not fill gaps with guesses.- For contract-layer changes, give before -> after, the reason, and a one-line note on the impact for existing users.- Name at most 3 files a human should read first, with line numbers.- If you deferred any judgment call, write it down along with why you hesitated. Do not omit it.
I added that last line after realizing that silently swallowed hesitation is the worse outcome. The places where an agent deferred a decision are exactly the places I most want to check the next morning.
## Re-entry digest (since last review point)- Irreversible / external: none (push not run, no deploy)- Contract: changed Article price in pricing.ts 250 -> 280 (reason: spec revision. backward compat: no effect on existing purchases)- Look first: src/config/pricing.ts L42 / src/lib/premium.ts L18- Deferred: one test still references the old price; left unmodified on purpose- Noise: formatting only across 31 files (nothing to check)
With this digest, re-entry can begin from the few files listed under "look first." The big picture comes from the risk classes, the detail from following those lines — a two-tier approach.
Facts From the Machine, Reasons From the Agent
If the agent writes the entire digest, then file counts and filenames become generated content too. A single number being off costs you trust in the whole review, so I eventually split the work: git supplies the facts, and the agent supplies only the reasoning.
#!/usr/bin/env bash# scripts/digest.sh — assemble the factual halfset -euo pipefailAGENT="${1:-antigravity}"BASE=$(git rev-parse -q --verify "reviewed/${AGENT}" || git merge-base origin/main HEAD)echo "## Re-entry digest (${AGENT} / ${BASE:0:7}..HEAD)"echoecho "### Change size"git diff --shortstat "${BASE}..HEAD"echoecho "### By risk class"node scripts/classify-changes.mjs "$BASE"echoecho "### External side effects (the real thing)"echo "- vs origin/main: $(git rev-list --count "origin/main..HEAD") commits ahead"echo "- tags on remote: $(git ls-remote --tags origin | grep -c 'refs/tags' || true)"
Once the skeleton of the output is fixed on the machine side, what you ask of the agent is filling blanks. Reasons, backward compatibility, and deferred calls all need human language, so those stay with the agent.
After the split, I stopped spending time doubting numbers in the digest. Not having to decide whether to trust a generated artifact every single morning turned out to matter more than I expected.
How Far to Trust the Self-Report
This is the most important trap. The agent's self-reported digest is good enough for classifying internal and noise changes. But the irreversible / external class alone is the one you must not take at face value.
Even when the agent writes "push not run," I always verify git log origin/main and the actual deploy history with my own eyes. The reason is simple: the operations hardest to take back cause the most damage when misclassified. Use the digest as an entry map, and verify only the irreversible operations against the real thing. Drawing that line lets you trust the digest while preventing accidents.
Make Verification a 30-Second Job
If verification runs on "remember to be careful," you will skip it on a busy morning. The way to make it unskippable is to reduce it to a single command.
#!/usr/bin/env bash# scripts/verify-irreversible.sh — check only the irreversible side against realityset -euo pipefailgit fetch --quiet origin --tags --pruneecho "== recent commits on remote main (anything not mine?) =="git log origin/main --since="18 hours ago" --pretty=" %h %an %s"echo "== divergence between local and remote =="git rev-list --left-right --count origin/main...HEAD | awk '{print " behind="$1" ahead="$2}'echo "== recent deploy runs =="gh run list --limit 5 --json displayTitle,status,conclusion,createdAt \ --jq '.[] | " \(.createdAt[0:16]) \(.conclusion // .status) \(.displayTitle)"'echo "== leftover untracked artifacts =="git status --porcelain --untracked-files=all | head -20
What you are reading here is the remote and the CI record, not the agent's words. --left-right --count gives you ahead and behind at a glance; a nonzero behind means you are reviewing against a stale local state.
The untracked-files check came later. I once found fragments of generated secrets and temporary files sitting in my working tree, and those show up in neither the diff nor the history.
The Days the Classifier Was Wrong
Rule-based classification does get things wrong. Two cases from my own repositories are worth recording.
The first was keeping lockfiles in the noise class. I skipped them the way I skip formatting, and a dependency had moved up a major version. The build passed, so I only noticed days later. Lockfiles have lived in the contract class ever since. Line count and importance have nothing to do with each other — that is where I learned it.
The second was a deletion hidden behind renames. Buried in a bulk rename of generated files, one route had disappeared. An R in --name-status looks harmless, so the eye slides right past it. Deletions and renames now get promoted a class mechanically. The UP map in the script above is a product of that failure.
What both have in common is how quietly they failed. Misclassification does not raise a warning. That is why, whenever I fix a rule, I leave a comment naming the failure it came from.
What Actually Got Shorter
Since impressions are unreliable, I timed two weeks of morning reviews and compared. This is one person's log against unattended runs of comparable size — not a rigorous measurement, but enough to read the trend.
Task
Before
After
Why it changed
Deciding where to start
~8 min
under 1 min
Changes arrive in class order
Reading the diff
~25 min
~9 min
The noise class never gets opened
Verifying irreversible ops
Skipped some days
~40 sec
Collapsed into one command
Scrolling back through chat logs
~12 min
Rarely needed
Deferred calls arrive up front
The time saved mattered less than the fact that there are no longer days when verification gets skipped. When the range you can safely skip is fixed, a person reads faster and more accurately.
There is added work too: maintaining the rules. Every time the project structure shifts, a few lines need updating, and neglecting them lets classification decay quietly. I set aside time once a month to read the rule table from the top.
As a next step, add one line to the end of the task you currently run unattended, requiring it to write the four digest items (irreversible, contract, look first, deferred). The classification script can come later. Your next-morning re-entry will stop stalling on where to begin.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.