◉ANTIGRAVITY LABJP
Articles/Agents & Manager
◈ Agents & Manager/2026-06-29Advanced

Spotting Agents That Are Alive but Stuck — Designing a Progress Heartbeat and Watchdog

The process is alive but the work isn't moving — the nastiest state for a background agent. Here is how to switch from liveness to progress monitoring to detect it, and how to stop it safely, with working code.

background agents3Antigravity376watchdogtimeout designoperations33

✦ Premium Article

In the morning, the dashboard was still green. The process was running, it was eating a little CPU, the last log line read "querying the model" — and yet not a single character had advanced in six hours.

As an indie developer, I run an article-generation agent overnight. That night an external API never responded, never even entered retry, and just kept waiting. The process was genuinely alive. So the monitoring that watched "is it alive" kept returning green the whole time. The mistake was treating liveness and progress as the same thing.

"Alive" and "progressing" are different questions

Out of server-monitoring habit, we reach for "does the process exist" and "does the port respond" as health signals. For short requests that is enough. But a background agent can spend minutes to tens of minutes on a single job, and throughout it can be "perfectly present while advancing nothing."

The classic stalls are an external call that never returns, an infinite reasoning loop in front of an unsolvable dependency, and a deadlock where it waits on itself. In all of them the process is alive. Liveness monitoring catches none of them. What you need is progress monitoring.

Kind of monitoringQuestion it answersCatches a stall?
Liveness (process exists, port responds)Is it running?No
Log output presentIs it talking?Fooled by endless "waiting…"
Progress (amount advanced)Did work move forward?Yes

Judging by whether logs are flowing is also dangerous. Code that keeps printing "waiting…" is talking but not advancing. Watch the amount of forward motion, not the chattiness.

Emitting progress as a marker

To watch by progress, the agent itself must leave "I got this far" in an externally observable form. Write a monotonically increasing step number and the time it last advanced to a shared location.

# progress.py — stamp a progress marker into a shared file
import json, os, tempfile, time
 
def mark_progress(path: str, step: int, note: str = "") -> None:
    """step is monotonic. Call only when you advanced (not while waiting)."""
    payload = {"step": step, "note": note, "ts": int(time.time())}
    fd, tmp = tempfile.mkstemp(dir=os.path.dirname(path) or ".")
    with os.fdopen(fd, "w") as f:
        json.dump(payload, f)
    os.replace(tmp, path)        # atomic swap

The discipline that matters: call mark_progress only when you advanced. If you design the heartbeat to "always fire on a timer," the pulse continues even while stalled, and you are back to liveness monitoring. A pulse must mean "moved forward," not "still alive."

In the agent body, advance the step at meaningful boundaries.

mark_progress(P, 1, "collection done")
# ... external API call (this is where it tends to freeze) ...
mark_progress(P, 2, "draft generated")
# ... formatting ...
mark_progress(P, 3, "verification passed")
✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦Understand that 'is it alive' and 'is it progressing' are entirely different questions, and why liveness monitoring alone misses a silent stall
✦A complete implementation from progress heartbeat to no-progress timeout to safe termination (a Python watchdog plus a bash integration). Drop it straight into your own job
✦From the real experience of running an agent overnight as an indie developer and finding it frozen silently until morning, a clear rule for where to place progress markers
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

◈ Agents & Manager2026-09-06
With scheduled agents, I now look for runs that never happened before I look for failures
My run ledger showed a 100% success rate while the evening slot had not fired for two weeks. Here is the reconciliation I now run against an expected-fire table, with measured notes on cron expansion, exit codes, and how the ledger itself gets written.
◈ Agents & Manager2026-04-26
Designing Antigravity Agent Traces That Tell You Why It Failed — Observability in Practice
Run Antigravity agents long enough and unreadable failure logs pile up fast. This piece walks span structure, attribute design, failure tagging, dashboards, cost visibility, and retry policy — backed by six months of production metrics — so you can cut post-incident debugging time in half.
◈ Agents & Manager2026-09-09
When an Antigravity Scheduled Task Never Runs: Check sidecar.json, Then enabled, Then the Logs
A checking order for Antigravity 2.0 Scheduled Tasks that quietly do nothing: where sidecar.json lives, how the directory name becomes the ID, the enabled flag in config.json, and the logs under sidecar_data.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links