◉ANTIGRAVITY LABJP
Articles/Agents & Manager
◈ Agents & Manager/2026-04-13Advanced

Building a Coding Agent System with Gemma 4 × Antigravity — A Complete Implementation Guide for Code Review, Test Generation, and Refactoring

A hands-on guide to building a 3-agent collaborative system using Gemma 4 and Antigravity AgentKit 2.0, covering code review, automated test generation, and refactoring suggestions. Includes production-quality code and pitfall solutions.

gemma412antigravity461agentkit13ai-agent18code-review11testing17refactoring8

Coding agents have a reputation for working great in demos but falling apart on real codebases. The usual culprit is a design that delegates everything to a single model call. Code review, test generation, and refactoring suggestions each demand different context strategies and different recovery paths when something goes wrong.

Gemma 4 has become a practical choice for solo developers because it runs locally and offers low-cost API calls. You can experiment without watching your bill, then switch to the cloud API when you're ready for production. This guide walks through building a system using Antigravity's AgentKit 2.0 that assigns each of those three responsibilities to a dedicated agent and coordinates them together.

What Makes a Coding Agent Actually Useful

Before Gemma 4, running a coding agent in production as a solo developer meant facing a cost wall. High-capability models like GPT-4 or Gemini 2.5 Pro understand code well, but continuously processing a large codebase gets expensive fast.

On my own setup, a single review of a ~500-line diff landed at roughly one twentieth of what the same job cost through Gemini 2.5 Pro. The absolute numbers matter less than the change in order of magnitude. When 50 PRs a day keeps the monthly bill in single-digit dollars, you stop weighing cost every time you consider running a review. Removing that decision was the real win for me.

But choosing purely on cost leads to disappointment. Here is what building this system actually revealed about where the model helps and where it does not:

Where Gemma 4 performs well: code style and best practice reviews, test code generation that mirrors existing test patterns, and focused refactoring suggestions at the function or class level.

Where Gemma 4 struggles: cross-file architectural issues, independently judging what deserves test coverage in an unfamiliar codebase, and high-abstraction architecture reviews.

This one gets misread often, so let me be precise. The reason is not a small context window. The Gemma 4 model card puts 12B, 26B A4B, and 31B at 256K tokens, with E2B and E4B at 128K. Dozens of files physically fit.

The question is not whether the code fits but whether the model can still recall it. On the same card's long-context benchmark (MRCR v2, eight needles in 128K), 31B scores 66.4% and 26B A4B scores 44.1%. Design for losing close to half of it. That is why the architecture below narrows context on the way in rather than handing over everything and hoping for good judgment.

Given these characteristics, "one agent using Gemma 4 for everything" is the wrong architecture. A "three-agent setup where each agent only works in Gemma 4's strong zone" is the right one.

System Architecture — Three-Agent Collaboration

The system built in this guide looks like this:

Manager Agent (Antigravity AgentKit 2.0)
  ├── Code Review Agent    ← Gemma 4 (diff analysis + style checks)
  ├── Test Generation Agent ← Gemma 4 (pattern-guided test generation)
  └── Refactoring Agent    ← Gemma 4 (function/class-scoped suggestions)

Why the Manager Agent pattern

AgentKit 2.0's Manager Agent acts as a coordinator — it receives a task, delegates to the right subagent, and collects results. The key benefit is that each subagent can fail and retry independently. If the Code Review Agent fails, Test Generation still proceeds.

With a single-agent design, one failure stops everything. The Manager Agent pattern separates "which agent handles what" from "how each agent does its job."

Context strategy by agent

Each agent deliberately constrains what it sends to Gemma 4:

  • Code Review Agent: git diff only (before and after the change)
  • Test Generation Agent: target file + existing test file + test framework config
  • Refactoring Agent: target function or class only (not the entire file)

This maximizes Gemma 4's limited context window by following a strict principle: send only what's necessary.

Environment Setup

Initialize an AgentKit 2.0 project inside Antigravity:

npm create agentkit@latest my-coding-agent
cd my-coding-agent
npm install
 
# Add to .env.local:
# GEMMA_API_KEY=YOUR_GEMMA_API_KEY
# GEMMA_MODEL=gemma-4-26b-a4b-it

Project structure:

my-coding-agent/
  ├── src/
  │   ├── agents/
  │   │   ├── manager.ts
  │   │   ├── code-review.ts
  │   │   ├── test-gen.ts
  │   │   └── refactor.ts
  │   ├── tools/
  │   │   ├── git-tools.ts
  │   │   ├── file-tools.ts
  │   │   └── github-tools.ts
  │   └── index.ts
  ├── .env.local
  └── package.json

Build the Gemma 4 client first, with error handling wired in from the start. Retrofitting error handling into an agent system is painful because state management between agents gets complicated quickly.

Before that, clear the trap that costs people the most time. The Gemini API serves exactly two Gemma 4 model IDs: gemma-4-31b-it and gemma-4-26b-a4b-it (Run Gemma with the Gemini API). Carry a Gemma 3 habit over and reach for a 27B-style ID and the call simply fails, because that size does not exist in the Gemma 4 family. The error only echoes the model name back, so discovering it after writing all three agents turns into a long detour. Verify one round trip before you build anything else.

The default here is gemma-4-26b-a4b-it. It is a MoE model with 25.2B total parameters but only 3.8B active at inference, so it answers noticeably faster than 31B Dense while scoring 77.1% on LiveCodeBench v6 against 31B's 80.0%. In day-to-day review that gap is hard to feel; the latency gap is not. I switch to 31B only when accuracy outranks speed, such as design-level reviews.

One note on the SDK. The JavaScript package @google/generative-ai reached end of support on August 31, 2025, and the current library is @google/genai. The code below assumes the newer SDK.

// src/lib/gemma-client.ts
import { GoogleGenAI } from "@google/genai";
 
const ai = new GoogleGenAI({ apiKey: process.env.GEMMA_API_KEY });
 
export interface GemmaCallOptions {
  model?: string;
  maxOutputTokens?: number;
  temperature?: number;
  retries?: number;
}
 
export async function callGemma(
  systemPrompt: string,
  userMessage: string,
  options: GemmaCallOptions = {}
): Promise<string> {
  const {
    model = "gemma-4-26b-a4b-it",
    maxOutputTokens = 2048,
    temperature = 0.2, // Low temperature for reproducible coding tasks
    retries = 3,
  } = options;
 
  for (let attempt = 1; attempt <= retries; attempt++) {
    try {
      const response = await ai.models.generateContent({
        model,
        contents: userMessage,
        config: {
          systemInstruction: systemPrompt,
          maxOutputTokens,
          temperature,
        },
      });
 
      const text = response.text;
 
      if (!text || text.trim().length === 0) {
        throw new Error("Empty response from Gemma 4");
      }
 
      return text;
    } catch (error: unknown) {
      const isLastAttempt = attempt === retries;
      const errorMessage = error instanceof Error ? error.message : String(error);
 
      if (isLastAttempt) {
        throw new Error(
          `Gemma 4 call failed after ${retries} attempts: ${errorMessage}`
        );
      }
 
      // Always wait, not just on 429. Retrying instantly just burns
      // all three attempts on the same failure.
      const isRateLimited =
        errorMessage.includes("429") || errorMessage.includes("rate limit");
      const waitMs = isRateLimited ? 2 ** attempt * 1000 : 500;
      await new Promise((resolve) => setTimeout(resolve, waitMs));
    }
  }
 
  throw new Error("Unreachable");
}

The temperature: 0.2 choice deserves explanation. Code review needs to be reproducible — reviewing the same diff twice and getting contradictory feedback erodes developer trust quickly. 0.2 keeps quality intact without making the wording feel mechanical.

This is a deliberate departure from the official guidance. The model card recommends temperature=1.0, top_p=0.95, and top_k=64 across all use cases, and the published benchmark scores come from that configuration. If reproducing benchmark-level accuracy matters more to you, go back to the recommended values. I chose consistency of output over peak accuracy. Neither answer is wrong — decide which one your team cares about before the first disagreement about a review comment, not after. I went with reproducibility because unstable review output costs more in practice than the accuracy I gave up.

Gemma 4 also introduces a thinking mode. Through the Gemini API you toggle it with thinkingConfig.thinkingLevel, set to "high" or "minimal". Leaving it on "high" for routine style checks mostly buys you latency rather than better findings, while refactoring suggestions genuinely benefit from it. Set it per agent rather than globally.

Code Review Agent Implementation

The core contract of the Code Review Agent is simple: receive a diff, return structured feedback as JSON. Unstructured text output creates downstream parsing nightmares.

// src/agents/code-review.ts
import { callGemma } from "../lib/gemma-client";
 
export interface ReviewComment {
  file: string;
  line: number;
  severity: "error" | "warning" | "suggestion";
  category: "security" | "performance" | "style" | "logic" | "test-coverage";
  message: string;
  suggestion?: string;
}
 
export interface ReviewResult {
  comments: ReviewComment[];
  summary: string;
  approvalStatus: "approved" | "changes-requested" | "comment";
}
 
const CODE_REVIEW_SYSTEM_PROMPT = `You are an expert code reviewer. Analyze the provided git diff and return feedback as valid JSON only.
 
Rules:
- Focus only on the changed lines in the diff
- Do NOT comment on unchanged code
- Return JSON matching the ReviewResult type exactly
- If no issues found, return empty comments array with "approved" status
- severity "error" = blocks merge, "warning" = should fix, "suggestion" = optional
- Return ONLY valid JSON. Do NOT wrap in markdown code fences.`;
 
export async function runCodeReviewAgent(
  diff: string,
  prTitle: string,
  prDescription: string
): Promise<ReviewResult> {
  const MAX_DIFF_CHARS = 8000;
  let processedDiff = diff;
 
  if (diff.length > MAX_DIFF_CHARS) {
    // Strip unchanged context lines — they don't improve Gemma 4's accuracy
    processedDiff = diff
      .split("\n")
      .filter(
        (line) =>
          line.startsWith("+") ||
          line.startsWith("-") ||
          line.startsWith("@@") ||
          line.startsWith("diff")
      )
      .join("\n")
      .slice(0, MAX_DIFF_CHARS);
 
    processedDiff += "\n[... diff truncated for context limit ...]";
  }
 
  const userMessage = `PR Title: ${prTitle}
PR Description: ${prDescription}
 
Git Diff:
\`\`\`diff
${processedDiff}
\`\`\`
 
Review this code change and return JSON only.`;
 
  const rawResponse = await callGemma(CODE_REVIEW_SYSTEM_PROMPT, userMessage, {
    temperature: 0.1,
    maxOutputTokens: 1024,
  });
 
  try {
    // Gemma 4 sometimes wraps JSON in code fences despite instructions
    const cleaned = rawResponse
      .replace(/^```json\n?/, "")
      .replace(/\n?```$/, "")
      .trim();
 
    return JSON.parse(cleaned) as ReviewResult;
  } catch {
    // Return a graceful fallback rather than crashing the whole pipeline
    console.error(
      "Failed to parse code review response:",
      rawResponse.slice(0, 200)
    );
    return {
      comments: [],
      summary:
        "Review completed but response parsing failed. Manual review recommended.",
      approvalStatus: "comment",
    };
  }
}

Two design decisions worth highlighting: stripping unchanged context lines before sending to Gemma 4, and returning a graceful fallback on JSON parse failure. The first keeps the model focused on what actually changed. The second means a broken review response doesn't take down test generation or refactoring for the same PR.

Test Generation Agent Implementation

The test generation agent uses a fundamentally different context strategy. The most important variable isn't the source file — it's the existing test file.

Every team has implicit conventions for tests: how describe and it blocks are nested, how mocks are written, what level of assertion granularity is expected. The most effective way to get Gemma 4 to follow those conventions is to show them by example in the user message — not the system prompt.

// src/agents/test-gen.ts
import { callGemma } from "../lib/gemma-client";
 
export interface TestGenerationResult {
  testCode: string;
  testFilePath: string;
  testFramework: "vitest" | "jest" | "unknown";
  coveredScenarios: string[];
  missingScenarios: string[]; // Scenarios Gemma 4 identified but chose not to generate
}
 
const TEST_GEN_SYSTEM_PROMPT = `You are an expert test engineer specializing in TypeScript/JavaScript testing.
 
Your task: Generate test code following the EXACT style of the existing tests.
 
Rules:
- Match the existing test structure precisely
- Use the same mocking approach as existing tests
- Cover: happy path, edge cases, error cases
- Return ONLY valid JSON. Do NOT wrap in markdown code fences.
- The test code must be ready to run without modification`;
 
export async function runTestGenerationAgent(params: {
  sourceFile: string;
  sourcePath: string;
  existingTestFile?: string;
  existingTestPath?: string;
  testFramework: string;
}): Promise<TestGenerationResult> {
  const { sourceFile, sourcePath, existingTestFile, existingTestPath, testFramework } = params;
 
  const existingTestSection = existingTestFile
    ? `## Existing Test Example (FOLLOW THIS STYLE EXACTLY)
File: ${existingTestPath}
\`\`\`typescript
${existingTestFile.slice(0, 3000)}${existingTestFile.length > 3000 ? "\n// ... (truncated)" : ""}
\`\`\``
    : `## No existing tests found. Use ${testFramework} best practices.`;
 
  const userMessage = `${existingTestSection}
 
## Source File to Test
File: ${sourcePath}
\`\`\`typescript
${sourceFile}
\`\`\`
 
Generate comprehensive tests. Return JSON only:
{"testCode":"...","testFilePath":"...","testFramework":"vitest|jest|unknown","coveredScenarios":["..."],"missingScenarios":["..."]}`;
 
  const rawResponse = await callGemma(TEST_GEN_SYSTEM_PROMPT, userMessage, {
    maxOutputTokens: 2048,
    temperature: 0.1,
  });
 
  try {
    const cleaned = rawResponse
      .replace(/^```json\n?/, "")
      .replace(/\n?```$/, "")
      .trim();
 
    const result = JSON.parse(cleaned) as TestGenerationResult;
 
    if (!result.testCode || result.testCode.trim().length < 50) {
      throw new Error("Generated test code is too short or empty");
    }
 
    return result;
  } catch (error) {
    const errorMessage = error instanceof Error ? error.message : String(error);
    throw new Error(`Test generation failed: ${errorMessage}`);
  }
}

The missingScenarios field is a deliberate design choice. By explicitly asking Gemma 4 to surface what it chose not to test, you give reviewers a concrete checklist of gaps. In practice, async error scenarios and null input boundary cases appear in this field frequently — exactly the kind of thing that bites you in production.

Refactoring Agent Implementation

The refactoring agent takes a different approach: it returns not just the refactored code, but the reason behind the change. This lets engineers make the call themselves rather than blindly applying AI suggestions.

// src/agents/refactor.ts
import { callGemma } from "../lib/gemma-client";
 
export interface RefactoringProposal {
  original: string;
  refactored: string;
  rationale: string;
  impactLevel: "low" | "medium" | "high";
  breakingChange: boolean;
  estimatedEffort: "minutes" | "hours" | "days";
  relatedPatterns: string[];
}
 
const REFACTOR_SYSTEM_PROMPT = `You are a senior software architect specializing in code refactoring.
 
Analyze the code and return a single, highest-priority refactoring proposal as JSON.
 
Priority order:
1. Security vulnerabilities
2. Performance bottlenecks (N+1, memory leaks, unnecessary loops)
3. Readability and maintainability
4. Design pattern improvements
 
Return ONLY valid JSON. Do NOT include markdown code fences in JSON values.`;
 
export async function runRefactoringAgent(params: {
  targetCode: string;
  language: string;
  context?: string;
}): Promise<RefactoringProposal | null> {
  const { targetCode, language, context } = params;
 
  if (targetCode.trim().split("\n").length < 5) {
    return null; // Not worth refactoring
  }
 
  const userMessage = `Language: ${language}
${context ? `Context: ${context}\n` : ""}
Code to refactor:
\`\`\`${language}
${targetCode}
\`\`\`
 
Return the highest-priority refactoring as JSON:
{"original":"...","refactored":"...","rationale":"...","impactLevel":"low|medium|high","breakingChange":false,"estimatedEffort":"minutes|hours|days","relatedPatterns":["..."]}`;
 
  const rawResponse = await callGemma(REFACTOR_SYSTEM_PROMPT, userMessage, {
    temperature: 0.15,
    maxOutputTokens: 1500,
  });
 
  try {
    const cleaned = rawResponse
      .replace(/^```json\n?/, "")
      .replace(/\n?```$/, "")
      .trim();
 
    return JSON.parse(cleaned) as RefactoringProposal;
  } catch {
    return null; // Non-fatal — let Manager Agent handle the absence
  }
}

Returning a single highest-priority proposal is intentional. Multiple suggestions create decision overhead: "which one do I fix first?" One clear priority eliminates that friction and makes the agent's output immediately actionable.

Manager Agent: Tying It All Together

// src/agents/manager.ts
import { runCodeReviewAgent, ReviewResult } from "./code-review";
import { runTestGenerationAgent, TestGenerationResult } from "./test-gen";
import { runRefactoringAgent, RefactoringProposal } from "./refactor";
 
export interface CodingAgentInput {
  diff: string;
  prTitle: string;
  prDescription: string;
  changedFiles: Array<{
    path: string;
    content: string;
    testFilePath?: string;
    testFileContent?: string;
    language: string;
  }>;
}
 
export interface CodingAgentOutput {
  review: ReviewResult | null;
  tests: TestGenerationResult[];
  refactoring: RefactoringProposal[];
  processingTime: number;
  errors: string[];
}
 
export async function runCodingAgentSystem(
  input: CodingAgentInput
): Promise<CodingAgentOutput> {
  const startTime = Date.now();
  const errors: string[] = [];
  const tests: TestGenerationResult[] = [];
  const refactoring: RefactoringProposal[] = [];
 
  // Fire code review immediately — it only needs the diff, not individual files
  const reviewPromise = runCodeReviewAgent(
    input.diff,
    input.prTitle,
    input.prDescription
  ).catch((err: Error) => {
    errors.push(`Code review failed: ${err.message}`);
    return null;
  });
 
  // Process each file in parallel
  const fileProcessingPromises = input.changedFiles.map(async (file) => {
    if (["ts", "tsx", "js", "jsx"].some((ext) => file.path.endsWith(`.${ext}`))) {
      try {
        const testResult = await runTestGenerationAgent({
          sourceFile: file.content,
          sourcePath: file.path,
          existingTestFile: file.testFileContent,
          existingTestPath: file.testFilePath,
          testFramework: detectTestFramework(file.testFileContent),
        });
        tests.push(testResult);
      } catch (err: unknown) {
        const errorMessage = err instanceof Error ? err.message : String(err);
        errors.push(`Test generation failed for ${file.path}: ${errorMessage}`);
      }
    }
 
    try {
      const refactorResult = await runRefactoringAgent({
        targetCode: file.content,
        language: file.language,
      });
      if (refactorResult) {
        refactoring.push(refactorResult);
      }
    } catch (err: unknown) {
      const errorMessage = err instanceof Error ? err.message : String(err);
      errors.push(`Refactoring analysis failed for ${file.path}: ${errorMessage}`);
    }
  });
 
  const [review] = await Promise.all([
    reviewPromise,
    Promise.all(fileProcessingPromises),
  ]);
 
  return {
    review,
    tests,
    refactoring,
    processingTime: Date.now() - startTime,
    errors,
  };
}
 
function detectTestFramework(testFileContent?: string): string {
  if (!testFileContent) return "vitest";
  if (testFileContent.includes("from 'vitest'") || testFileContent.includes('from "vitest"')) return "vitest";
  if (testFileContent.includes("from 'jest'") || testFileContent.includes("@jest/globals")) return "jest";
  return "vitest";
}

Running code review in parallel with file processing cuts total execution time by roughly 60% compared to sequential processing. Both operations are independent — the review only needs the diff, and file processing doesn't depend on the review result.

Deploying to Cloudflare Workers

Antigravity's Cloudflare Workers support lets you expose the coding agent as a webhook API, callable from GitHub Actions or any CI/CD pipeline.

// src/index.ts
import { runCodingAgentSystem, CodingAgentInput } from "./agents/manager";
 
interface Env {
  WEBHOOK_TOKEN: string;
  GEMMA_API_KEY: string;
}
 
export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    if (request.method !== "POST") {
      return new Response("Method not allowed", { status: 405 });
    }
 
    const authHeader = request.headers.get("Authorization");
    if (authHeader !== `Bearer ${env.WEBHOOK_TOKEN}`) {
      return new Response("Unauthorized", { status: 401 });
    }
 
    let input: CodingAgentInput;
    try {
      input = (await request.json()) as CodingAgentInput;
    } catch {
      return new Response("Invalid JSON body", { status: 400 });
    }
 
    // Race against Cloudflare's 30-second timeout
    const timeoutMs = 25000;
    const agentPromise = runCodingAgentSystem(input);
    const timeoutPromise = new Promise<never>((_, reject) =>
      setTimeout(() => reject(new Error("Processing timeout")), timeoutMs)
    );
 
    try {
      const result = await Promise.race([agentPromise, timeoutPromise]);
      return new Response(JSON.stringify(result), {
        headers: { "Content-Type": "application/json" },
      });
    } catch (error: unknown) {
      const errorMessage = error instanceof Error ? error.message : String(error);
      return new Response(JSON.stringify({ error: errorMessage }), {
        status: 500,
        headers: { "Content-Type": "application/json" },
      });
    }
  },
};

Pre-filtering to cut API costs

About 30% of commits in practice don't warrant AI review — documentation updates, config changes, type definition files. Filtering these out before making any API calls keeps costs down significantly.

function shouldSkipReview(diff: string): boolean {
  const nonCodeExtensions = [".md", ".txt", ".json", ".yml", ".yaml", ".toml"];
  const changedFiles = diff
    .split("\n")
    .filter((line) => line.startsWith("diff --git"))
    .map((line) => line.split(" b/")[1]);
 
  return changedFiles.every((file) =>
    nonCodeExtensions.some((ext) => file?.endsWith(ext))
  );
}

Five Production Pitfalls and How to Avoid Them

Pitfall 1: JSON wrapped in code fences

Gemma 4 sometimes wraps JSON responses in markdown code fences despite explicit instructions not to. The fix is defensive parsing — strip code fence markers before parsing, and add "Do NOT wrap in markdown code fences" to your system prompt.

Pitfall 2: Attention degradation on long context

Like most LLMs, Gemma 4's attention on later instructions weakens as context grows. In practice, functions in the second half of a file get reviewed less thoroughly. The long-context benchmark numbers quoted earlier are that effect made measurable.

Keep system prompts under 200 tokens and put critical instructions at the beginning of the user message.

Beyond that, measure whether it is actually happening in your setup. I keep one canary diff with a single obvious defect planted at the very end — an unused variable, or a flipped comparison operator — and check whether it gets flagged. Growing that diff step by step reveals the length at which your own configuration starts dropping the tail. Mine began missing things past roughly 800 changed lines, which is where I added per-file splitting.

Pitfall 3: Rate limits from parallel requests

Promise.all fires all requests simultaneously and often triggers rate limits. Use p-limit to cap concurrency:

import pLimit from "p-limit";
const limit = pLimit(3); // Max 3 concurrent requests
 
const fileProcessingPromises = input.changedFiles.map((file) =>
  limit(async () => {
    // file processing
  })
);

Pitfall 4: Cloudflare Workers' 30-second timeout

PRs touching more than 10 files will hit the timeout. The short-term fix (used in this article's code) is racing against a 25-second deadline and returning partial results. The proper long-term solution is moving to Durable Objects for async processing.

Pitfall 5: Silent failures killing observability

The first week in production, reviews were silently disappearing. The cause was catching all Gemma 4 errors and returning empty results with no logging. Always surface errors in the response body — the errors array in the output type exists exactly for this purpose.

One Month of Production Data

Running this system on a personal project for a month kept the API bill around $4. The equivalent with Gemini 2.5 Pro would have been roughly $80.

On accuracy: Gemma 4 misses things that Gemini 2.5 Pro would catch, particularly deep business logic issues and architectural problems. But for style violations, obvious bugs, unused variables, and test coverage gaps — routine problems that make up the bulk of real PR review work — Gemma 4's detection rate was comparable.

For an indie developer or a small team wiring AI review into CI/CD, a Gemma 4-based coding agent is a practical option — not as a replacement for human judgment, but as the thing that absorbs the routine findings so reviewers can spend their attention elsewhere.

As a next step, run just the Code Review Agent on gemma-4-26b-a4b-it against ten real diffs from your own repository. Learn where it misses before adding the test generation and refactoring agents; the ordering saves rework.

AgentKit 2.0 orchestration is covered in Antigravity AgentKit 2.0 — Production Multi-Agent Orchestration, and the guardrails that stop an agent from running away are in Antigravity Agent Safety Design — Guardrails and Runaway Prevention.

Share

Thank You for Reading

Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

If you found this article helpful, a small tip ($1.50) would mean a lot to us. Your support helps keep this site ad-free and covers server and hosting costs.

Related Articles

◈ Agents & Manager2026-05-03
Multi-Agent CI/CD Quality Gate — Automating PR Review, Testing, and Security with Antigravity × GitHub Actions
A complete implementation guide for building a multi-agent CI/CD quality gate using AgentKit 2.0 and GitHub Actions. Covers parallel code review, test generation, and security scanning agents — with production cost management and pitfall solutions.
◈ Agents & Manager2026-04-28
Versioning and A/B Testing Prompts for Production AI Agents in Antigravity
Hard-coding prompts in production turns improvement into guesswork. This guide walks through registry design, A/B traffic routing, statistical promotion, and rollback — with code you can ship inside your Antigravity environment.
◈ Agents & Manager2026-04-27
Building Self-Healing Antigravity Agents — Detection, Diagnosis, and Recovery in Production
A practical three-layer pattern for keeping Antigravity agents alive in production: signal-based detection, deterministic diagnosis, and graduated recovery — with full AgentKit 2.0 code and the production traps I learned the hard way.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links