Supervising Multiple Agents at Once on the Antigravity 2.0 Desktop: Screen Layout and Interruption Design
How I lay out the screen, score interruption order, and measure my own ceiling on parallel agents in Antigravity 2.0 — with the small scripts and the numbers I actually recorded.
As an indie developer running several apps in parallel, my desktop now runs three to four agents at once since Antigravity 2.0 was recast as an "agent control tower." Convenient as that is, for the first week I couldn't tell which one was waiting on what, ended up cycling through all of them in turn, and the parallelism gained me almost nothing.
The problem wasn't parallel execution itself. It was that my screen and my attention as the supervisor hadn't caught up to running them together. Coordinating multiple agents is less like writing code and more like air-traffic control over several things happening at once. Here I'll split that work into two parts: screen layout and interruption decisions.
What breaks first under parallelism is attention
When I ran a single agent, I just followed its output and there was nothing to agonize over. The moment I went to three, I tried to watch all of them equally and ended up watching each one only halfway.
What I realized is that a human can deeply follow essentially one agent at a time. The other two or three need to switch from "watch" to "watch when a cue arrives." In other words, designing parallel supervision turned out to be designing for fewer things to attend to.
Split the screen into "running," "needs decision," and "done"
So I started physically separating agent state into three zones.
Running: working autonomously right now, not awaiting human input. Don't look by default
Needs decision: stopped, waiting for confirmation or permission. Only this zone gets active attention
Done: finished. Review the results together
The crux of this three-way split is deciding deliberately not to look at the running zone. Peering at an agent mid-work only makes you anxious; it adds no basis for a decision. In my experience, pushing time spent on the running zone toward zero actually made my reactions to "needs decision" faster.
On the desktop, I prefix each agent with a state marker you can read at a glance.
Just putting these three states at the head of the task name makes it instantly clear, when you skim the list, that "the only thing I should look at now is WAIT." I find a text prefix easier to distinguish in peripheral vision than color coding.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦A 40-line board script that surfaces how long each WAIT has been sitting
✦A scoring formula that ranks interruptions by blocked work, volatility, and wait time
✦How to derive your parallelism ceiling from review-log median and p90, plus a 5-step recovery
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
After a few days on prefixes alone, one missing piece became obvious: how long a state has been sitting there.
A WAIT that appeared one minute ago and one that appeared twenty-five minutes ago call for very different urgency. But every row in the editor's list looks equally fresh, so age is invisible. I ended up keeping a single tab-separated state file and rendering it, sorted, in a corner of a second display.
Here's the file format. One line per agent: state, name, the moment it entered that state, and how many downstream tasks it blocks.
I chose this shape because an agent hook can append a line with a single printf, and I can hand-edit it without care. JSON was easy to corrupt when editing by hand, and too heavy for a scratch file whose only job is supervision.
The reader and sorter looks like this.
#!/usr/bin/env python3"""board.py — list running agents, WAIT first, with age.state.tsv, one line per agent: state<TAB>name<TAB>entered_state_at(ISO8601)<TAB>blocked_downstream"""from datetime import datetime, timezonefrom pathlib import PathSTATE_FILE = Path.home() / ".antigravity" / "state.tsv"ORDER = {"WAIT": 0, "DONE": 1, "RUN": 2}STALE_MIN = 10 # flag a WAIT older than thisdef load(path): rows = [] for line in path.read_text(encoding="utf-8").splitlines(): line = line.strip() if not line or line.startswith("#"): continue cols = (line.split("\t") + ["", "", "", "0"])[:4] state, name, since, blocking = cols rows.append({ "state": state.upper(), "name": name, "since": datetime.fromisoformat(since), "blocking": int(blocking or 0), }) return rowsdef minutes_since(ts): now = datetime.now(ts.tzinfo or timezone.utc) return int((now - ts).total_seconds() // 60)def main(): rows = load(STATE_FILE) for r in rows: r["age"] = minutes_since(r["since"]) rows.sort(key=lambda r: (ORDER.get(r["state"], 9), -r["blocking"], -r["age"])) for r in rows: mark = "!" if r["state"] == "WAIT" and r["age"] >= STALE_MIN else " " print(f"{mark}[{r['state']:<4}] {r['name']:<22} {r['age']:>4}m blocks {r['blocking']}")if __name__ == "__main__": main()
A word on why the sort key is what it is. It's a three-part tuple: (state rank, -blocked count, -age). The first part pins WAIT to the top permanently. The second lifts anything whose stall also stalls others. The third breaks ties toward whatever has been neglected longest. Arrival order is deliberately absent, because as the next section shows, arrival order is the most expensive ordering available.
The ! marker appears only on a WAIT older than ten minutes. It flags "you probably haven't noticed this," and the moment it shows up it's also telling you your own polling interval is too long.
To keep it live, one terminal is enough.
while true; do clear; python3 ~/bin/board.py; sleep 20; done
Twenty seconds is deliberate: at ten the numbers churn distractingly in peripheral vision, and at sixty the lag between a ! appearing and my noticing it became obvious. Since putting this panel next to my editor, the time from a WAIT appearing to my noticing it dropped from about four and a half minutes to roughly one, measured with a stopwatch.
Decide interruptions by the cost of making them wait
When several "needs decision" agents stop at the same time, which to handle first is the next problem. I prioritize not by arrival order but by the cost of keeping them waiting.
Two rules of thumb:
Is it blocking downstream work? If this one is stuck, the other two can't proceed — that's top priority
Is its context volatile? Something stopped mid-browser-operation goes stale if left, and tends to need a restart from scratch
Conversely, a one-off confirmation that affects nothing else can wait. If you dutifully handle things in arrival order, blocking tasks sit at the back of the queue and overall throughput drops. In fact, just switching priority from arrival order to "blocking first" raised the number of tasks I cleared per half-day by about 20%.
Score the interruption order
"Blocking first" alone still leaves you stuck when two items block the same amount of work. Deliberation is itself a cost paid by whoever is waiting, so I now compute the order from a flat formula.
priority = blocked_downstream × 10 + volatility + waiting_minutes × 0.2volatility: mid-browser-operation or short-lived auth = 5 mid-edit on local files = 3 files written, awaiting confirmation = 0
The coefficients encode three simple beliefs. Blocked downstream work halts two or three agents wholesale, so it gets an order of magnitude more weight. Volatility stands in for the probability of a redo, so it sits in the middle. Waiting minutes are a tiebreaker only — at 0.2 per minute, fifty minutes of neglect is worth one blocked downstream task.
Running the four items on my board through it:
Agent
Blocked
Volatility
Waiting (min)
Score
Order
migrate-db-schema
2
0
38
27.6
1
store-screenshot-upload
0
5
12
7.4
2
fetch-store-metadata
0
5
0
5.0
3
rename-test-fixtures
0
0
21
4.2
4
By arrival order, rename-test-fixtures would have come second. Scoring pushes store-screenshot-upload ahead instead — it's mid-upload to the store through a browser session, and leaving it means the session expires and the whole thing restarts. The formula and the felt cost agreed.
I don't treat this as an optimization, just as a device for removing hesitation. When two scores are close, the choice genuinely doesn't matter, so I stop thinking and work top-down. The less deliberation supervision requires, the lighter it gets.
An interruption etiquette that preserves state
When you interrupt a stopped agent, sloppily tacking on instructions breaks its context. What I follow is writing the interruption as an addendum to the current direction.
(bad interruption)"Actually, do it a different way"(good interruption)"Keep the direction so far, but for the auth part only, switch from session cookies to a token approach. Don't break the existing tests."
The former risks making the agent throw away the context built so far. The latter spells out what to keep and what to change, so you can nudge the direction while preserving in-progress state. Treat an interruption as a diff to append, not an overwrite of the plan, and you'll fail less.
Having a fixed skeleton keeps this intact even when I'm rushing. Mine is three lines.
KEEP: design decisions and tests made so far stay unchangedCHANGE: <target> from <current form> to <new form>LIMIT: <what must not break / what must not be touched>
Putting KEEP first is the part that matters. When I led with the change, agents reading from the top more than once headed toward rebuilding the whole thing. Simply reordering the three lines noticeably cut how often an interruption triggered a restart from zero.
The parallelism ceiling is set by your own review speed
Finally, the matter of how many to run. I once greedily pushed to six agents and it fell apart completely. Agents can run in parallel, but reviewing the finished work is still done by one human.
When review can't keep up, finished agents pile up holding their results, you're slow to notice the "needs decision" cues, and the whole thing jams. It was the same shape as an inspection step becoming the bottleneck on a production line.
These days I empirically set the ceiling at "the number whose completions I can fully review within 30 minutes." For me that's three agents, and four only when the tasks are all very light.
Measure the ceiling from your review log
Guessing at "what I can review in thirty minutes" means anchoring on a good day. So for about three weeks I logged how long each review actually took — one command right after finishing one.
# append one line after each reviewprintf '%s\t%s\t%s\n' "$(date +%F)" "review-pr-1284" "9" >> ~/review_log.tsv
The aggregation reports both median and p90. Median alone guarantees a jam on the day a heavy task shows up.
#!/usr/bin/env python3"""ceiling.py — derive a parallelism ceiling from a review log.review_log.tsv: date<TAB>task<TAB>review_minutes"""import statisticsimport sysfrom pathlib import PathBUDGET_MIN = 30 # minutes available for review per passdef main(path): mins = [] for line in Path(path).read_text(encoding="utf-8").splitlines(): cols = line.split("\t") if len(cols) < 3: continue try: mins.append(float(cols[2])) except ValueError: continue if not mins: print("review log is empty") return 1 med = statistics.median(mins) p90 = statistics.quantiles(mins, n=10)[8] if len(mins) >= 10 else max(mins) print(f"n={len(mins)} median={med:.1f}m p90={p90:.1f}m") print(f"ceiling: {int(BUDGET_MIN // med)} by median / {int(BUDGET_MIN // p90)} by p90") return 0if __name__ == "__main__": sys.exit(main(sys.argv[1] if len(sys.argv) > 1 else str(Path.home() / "review_log.tsv")))
Across my three weeks (142 entries) the median was 7.4 minutes and p90 was 16.2. That gives a ceiling of four by median and one by p90. What I settled on in practice was three — splitting the difference.
To read those numbers correctly: the p90 ceiling of one means "on a day of back-to-back heavy reviews, one is all you can carry," not that one should be the standing limit. I use the pair as a range — median on ordinary days, sliding toward p90 on days that include heavy review work.
Here's what I recorded per half-day while varying the launch count.
Agents running
Completed per half-day
Sent-back rate
Mean WAIT age
2
5.5
7%
2 min
3
7.0
9%
3 min
4
7.2
18%
9 min
6
4.8
31%
27 min
Going from three to four barely moved completions while doubling the send-back rate — the sign of review getting sloppy. At six, completions fell below the two-agent baseline. Parallelism doesn't plateau at the ceiling; past the ceiling it actively degrades.
A 5-step recovery when it jams
Even with a ceiling, weeks with lots of interruptions jam anyway — the run-up to an App Store submission is reliably one of them. When six agents seized up completely, this is the order I used to get back to normal. It took roughly fifty minutes.
Stop launching anything new. Without this first, nothing that follows can catch up
Review DONE oldest-first. The newest is tempting, but the older the item, the more time you spend reconstructing its context — neglect compounds
Answer all non-blocking WAITs in one pass. I reply "continue on the current plan; batch up anything that needs a decision and ask at the end," which cuts the number of stopped agents directly
Kill any RUN whose output hasn't changed in 30 minutes. Most are waiting on something external or circling; waiting longer changed nothing
Restart with one fewer agent. Returning to the old count right after a jam produces a second jam the same day
Of the five, step three did the most work. Four pending decisions on the board was itself what dulled my judgment, and three of them turned out to be "just keep going." The jam was caused by the count of pending decisions, not by the volume of work.
Something to try first
Next time you run two or more agents, start by prefixing the task names with [RUN], [WAIT], and [DONE]. That alone surfaces "the one thing to look at now." Once that feels natural, add state.tsv and board.py so that neglect becomes visible as a number.
I'm still searching for my own optimal parallelism count, but ever since I understood that supervision isn't about spreading attention but narrowing where it goes, running multiple agents got a lot easier. I hope it helps anyone else worn out in front of the control tower.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.