On a Thursday evening I rewrote the same request three times. Working as an indie developer, the time lost in a stall comes straight out of my own pocket. I asked Antigravity to rework an article listing page, looked at the diff, and rolled back to a checkpoint. I polished the wording and rolled back again. By the time the third attempt landed, I had spent longer than writing it myself would have taken.
The next morning I dumped my recent request log to CSV and read through it. Every request I had rolled back shared something — and it had nothing to do with how carefully the prompt was written. Every rolled-back request was heavy in either the number of files it touched or the number of decisions it contained.
Those are not matters of taste. They are quantities you can count before you press send. And if you can count them before you send, you can judge them before you send. This is the record of building that judge and running it against 80 of my own requests.
One metric: rework rate
"Output quality" is slippery. If a human grades it on a five-point scale, the score drifts with how tired that human is. So I narrowed it to a single binary.
Rework rate = the share of requests that were not accepted on the first pass, requiring a rollback, a substantial rewrite, or a course correction through follow-up instructions.
Recording it is trivial: label each request accepted or reworked. If you use Antigravity's checkpoints, the presence of a rollback is the label. I pulled mine from editor session logs, but a handwritten note works. What matters is that the label is binary. Binary labels correlate; five-point impressions do not.
Three weeks of requests came to 80. Overall rework rate: 41%. I had been redoing nearly half of everything I asked for. Seeing that number for the first time sat heavily in my chest.
Scoring a request before it is sent
The scorer reads the request text. Its only job is to stop oversized requests from leaving the keyboard, so it does no real parsing. It counts six factors.
| Factor | How it is counted | Weight | Why |
|---|---|---|---|
| Estimated files touched | Filenames, component names, enumerated "and also" clauses | ×2.0 | Past six files, later edits stop agreeing with earlier ones |
| Open decisions | "whatever makes sense", "the best approach", "you decide" | ×3.0 | The single strongest predictor of rework |
| Implementation plus tests | "and write tests", "with tests" | ×2.5 | Different modes of thought; one of them always gets sloppy |
| Bundled asks | "while you're at it", "also", "at the same time" | ×2.0 | Bundling hides added decisions |
| Standing rules inline | "never use any", "comments in Japanese" | ×1.0 | Pulls attention away from the actual ask |
| Non-self-contained work | "wire it up later", "for now", "just a stub" | ×1.5 | Intermediate artifacts create reconciliation work downstream |
The weights are not invented. I started every factor at 1.0, ran a logistic regression over the 80 labeled requests, kept only the ordering of the coefficients, and rounded to numbers I could hold in my head. Eighty samples do not justify precise coefficients. Preserving the order is enough.
# preflight_score.py — grade a request before it reaches the agent
import re
from dataclasses import dataclass
DECISION_MARKERS = [
"whatever makes sense", "the best approach", "you decide",
"as appropriate", "figure out", "up to you", "however you like",
]
TEST_MARKERS = ["and write tests", "with tests", "plus tests", "include tests"]
BUNDLE_MARKERS = ["while you're at it", "also", "at the same time", "and then"]
POLICY_MARKERS = ["never use any", "comments in Japanese", "naming convention", "formatter"]
DEFERRAL_MARKERS = ["wire it up later", "for now", "just a stub", "temporarily"]
FILE_PAT = re.compile(r"[\w./-]+\.(tsx?|jsx?|py|go|rs|mdx?|css)\b")
COMPONENT_PAT = re.compile(r"\b[A-Z][a-zA-Z0-9]{2,}\b")
WEIGHTS = {
"files": 2.0, "decisions": 3.0, "tests": 2.5,
"bundles": 2.0, "policies": 1.0, "deferrals": 1.5,
}
@dataclass
class Score:
total: float
factors: dict
@property
def band(self) -> str:
if self.total < 6:
return "GREEN"
if self.total < 12:
return "AMBER"
return "RED"
def _count(text: str, markers: list[str]) -> int:
lowered = text.lower()
return sum(lowered.count(m.lower()) for m in markers)
def score_request(text: str) -> Score:
files = len(set(FILE_PAT.findall(text))) + len(set(COMPONENT_PAT.findall(text)))
factors = {
# Up to three files is harmless. Penalize only the excess.
"files": max(0, files - 3),
"decisions": _count(text, DECISION_MARKERS),
"tests": _count(text, TEST_MARKERS),
"bundles": _count(text, BUNDLE_MARKERS),
"policies": _count(text, POLICY_MARKERS),
"deferrals": _count(text, DEFERRAL_MARKERS),
}
total = sum(WEIGHTS[k] * v for k, v in factors.items())
return Score(total=total, factors=factors)
if __name__ == "__main__":
import sys
s = score_request(sys.stdin.read())
print(f"score={s.total:.1f} band={s.band}")
for k, v in s.factors.items():
if v:
print(f" {k}: {v} (+{WEIGHTS[k] * v:.1f})")
sys.exit(0 if s.band != "RED" else 1)Penalizing only files beyond the third is a small implementation choice that turned out to matter. Scoring in direct proportion to file count put weight on tiny single-file requests and blurred the threshold. Flattening the harmless region to zero lets the threshold sit exactly where the danger starts.
The decision factor carries the heaviest weight because the data said so. "Add authentication" folds three judgments into four words: how sessions are managed, where the token lives, what happens on failure. The agent will not ask. It will quietly pick the most common option for each. When that choice collides with how the existing codebase does things, the diff breaks silently.