agy -p worked well on my machine, and the next thought was to put it in a container and let it run overnight. What worried me first wasn't agy itself. It was whether the caller would even notice if agy froze.
As an indie developer who runs a few small apps and sites on my own, I'd rather find a hang before it finds me. So instead of the real agy, I wrote a small shell script that hangs on purpose and built the caller-side safety nets against that. Every piece of code and every output below was run on my machine. This is not a record of running the real agy on Cloud Run. For that part I lean on a write-up by someone who did, and I only state what I could confirm against the official docs.
The order matters most. Decide where authentication comes from before the container exists. Then give the caller its own exit when the agent hangs. Prompt tuning comes after both.
Facts to settle first
A Zenn post documents calling agy from Docker and a Cloud Run job (Antigravity CLI (agy) from Docker / Cloud Run — pitfalls, in Japanese). Set against the official install page and a GitHub issue, this is what holds up.
| Topic | What I could confirm | Source |
|---|---|---|
| Headless auth | The official docs name a Gemini API key: set modelProvider to gemini in settings.json and pass GEMINI_API_KEY. The page says the environment variable alone has no effect | Official install page |
| Service account / ADC | In the headless-auth issue (#223) the reporter tried GOOGLE_APPLICATION_CREDENTIALS and others and still got the OAuth prompt. I found no official reply on the page when I read it | GitHub issue #223 |
| Why it worked locally | A cached browser login. The official docs also say a valid token profile in the OS keychain signs you in without a browser | Official install page / the Zenn post |
| Freezing on approval | With nobody to approve a tool call, the run waited until --print-timeout and returned partial output with exit code 0 | The Zenn post |
One caution. I did not find --print-timeout or --dangerously-skip-permissions on the install page I read. The Zenn post uses them, so check your own version with agy --help before relying on them. Versions were not pinned there either: 1.0.15 locally and 1.2.16 in the container, as the post reports.
Safety net 1: don't pipe the installer straight into bash
According to the Zenn post, the installer fetched with curl ... | bash sometimes came back as gzip without a Content-Encoding header, from the same URL. I did not reproduce that. What I can check is what bash does with gzip bytes, and whether the "save, test, then run" pattern behaves.
printf '#!/bin/bash\necho INSTALL_OK\n' > plain.sh
gzip -c plain.sh > gz.raw
cp plain.sh plain.raw
for f in plain.raw gz.raw; do
if gzip -t "$f" 2>/dev/null; then
gunzip -c "$f" > out.sh; echo "$f: gzip -> decompressed"
else
cp "$f" out.sh; echo "$f: plain"
fi
bash out.sh
doneOutput:
plain.raw: plain
INSTALL_OK
gz.raw: gzip -> decompressed
INSTALL_OKPiping the gzip file straight into bash made it try to read binary bytes as a command and end without a useful message. A build log that holds only one garbled "bash: line 1: ..." is slow to diagnose. Three lines of detection avoid that.
In a real Dockerfile I would create a non-root user, run the detection-guarded install, and finish with agy --version so a missing binary fails at build time. The Zenn post also notes that the installer edits ~/.bashrc, which CMD never reads, so PATH has to be set with ENV.
Safety net 2: give the caller a deadline and a closed stdin
Here is a fake agy that reproduces the approval hang. Without --dangerously-skip-permissions it waits, then prints a partial-output message and exits 0.
#!/bin/bash
# fake_agy.sh — a fake agy that hangs on tool approval
for a in "$@"; do [ "$a" = "--dangerously-skip-permissions" ] && SKIP=1; done
if [ -z "$SKIP" ]; then
echo "waiting for tool approval..."
sleep ${FAKE_HANG:-30}
echo "print timeout after 2m0s with turn in progress; returning partial output"
exit 0
fi
echo OK > "${FAKE_OUT:-result.txt}"
echo "done"; exit 0The caller is a Python function with three rules: stdin is DEVNULL so nothing blocks on input, the caller keeps its own deadline, and success is decided by whether the expected file exists, not by the exit code.
import os, subprocess, time, pathlib
def run_agent(cmd, workdir, expect_file, timeout_s, env=None):
"""Don't trust the exit code. Decide by deadline, closed stdin, and output."""
env = {**os.environ, **(env or {})}
for k in ("APP_PASSWORD", "SESSION_SECRET"): # don't hand unneeded secrets to the agent
env.pop(k, None)
proc = subprocess.Popen(cmd, cwd=workdir, stdin=subprocess.DEVNULL,
stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
text=True, errors="replace", env=env)
deadline = time.time() + timeout_s
killed = False
while proc.poll() is None:
time.sleep(0.2)
if time.time() > deadline:
proc.kill(); proc.wait(); killed = True; break
out = proc.stdout.read()
produced = pathlib.Path(workdir, expect_file).is_file()
return {"exit": proc.returncode, "killed": killed,
"produced": produced, "ok": produced and not killed,
"tail": out[-200:]}I ran four situations, shortening the deadline and varying how long the fake hangs.
| Situation | exit | killed | produced | ok |
|---|---|---|---|---|
| No flag, 3 s deadline, hangs 10 s | -9 (reaped after kill) | True | False | False |
| No flag, 10 s deadline, returns "partial output" after 2 s | 0 | False | False | False |
| Flag set, writes the output, exits | 0 | False | True | True |
| Flag set, writes the output under a different name | 0 | False | False | False |
Rows two and four are the point of this article. The exit code is 0 and ok is still False. A check that treats exit 0 as success counts both as wins and passes an empty directory downstream.
My first version of the function had no wait() after kill(), so the first row logged exit: None. Adding the reap so that -9 shows up was a fix I only noticed because I ran the fake.
Safety net 3: verify that the tools the agent uses actually work
The part of the Zenn post that stayed with me: the agent reported "done" while the log showed a custom evaluation script raising an exception on every run. The agent treated that as an environment problem, skipped the evaluation, and finished without saying so.
A fake agy can't reproduce this, so I'm reporting a rule I took from the write-up, not something I tested.
- Give the job a self-test mode that runs each tool the agent will call once, on its own, in addition to checking that agy answers
- Run the self-test with the same arguments as production (the post says a self-test that omitted the flag reproduced the approval hang)
- Decide success from the caller's inspection of the output, never from the agent's report
The expect_file check in safety net 2 is the smallest form of "the caller inspects". Where existence is not enough, content checks go there.
The first decision: where authentication comes from
Back to the opening point. The official headless route is a Gemini API key, and the environment variable alone is not enough; settings.json needs modelProvider set to gemini. I have not verified that against the real agy myself, so I'm quoting the docs and nothing more.
If you want a service account or ADC, issue #223 shows no official support at the time I read it. "It worked locally, so it will work in the container" may only mean a login is cached in your keychain. Before building the image, run agy -p once with every browser login discarded. It is the cheapest check there is.
What bridges "it worked on my machine" and "it keeps working unattended" is the caller's own inspection.
A good first step: open agy --help and confirm that --print-timeout and --dangerously-skip-permissions exist in your version. For how approvals themselves behave, diagnosing a command approval dialog that keeps reappearing is a useful companion.