◉ANTIGRAVITY LABJP
Articles/AI Tools
⚙ AI Tools/2026-03-30Advanced

Antigravity × Custom AI Chatbot — Notes from Building RAG, Function Calling, and Streaming UI

Build notes for a custom AI chatbot with RAG, Function Calling, and Streaming UI, including the token-estimation bug that silently discarded 75% of the corpus.

antigravity459ai-chatbotrag8function-calling5streaming-uivercel-ai-sdkcloudflare-workers9vector-search3production71

✦ Premium Article

Setup and context — the recall problem that started this

When I wired a small RAG chatbot into the support flow for one of my wallpaper apps, the answers were mediocre for the first few weeks. Roughly 1,200 FAQ entries were sitting in the index, yet asking "how do I cancel" would not surface the cancellation steps in the top results. The embedding model and the vector database were both configured straight from the official samples. Nothing threw an error.

It took me an embarrassingly long time to find the cause. It was one line in the chunker: Math.ceil(text.length / 4). That estimate is off by roughly 3.9x for Japanese text. Chunks I believed were capped at 512 tokens were actually running past 2,000 — and the embedding model's input limit is 512. About three quarters of every document I fed in was being discarded silently, with no error and no warning.

What follows is the full build, including that detour: RAG, Function Calling, and Streaming UI combined into one working stack, with the code you need to run it.

  • RAG (Retrieval-Augmented Generation): Searches your own documents and databases to improve answer accuracy and reduce hallucinations
  • Function Calling: Dynamically connects to external APIs and databases to fetch real-time information or perform actions
  • Streaming UI: Displays token-by-token responses in real time, dramatically improving perceived response speed

This article is aimed at engineers with experience building AI applications, assuming familiarity with TypeScript, Next.js, and vector databases. If you'd like to learn RAG fundamentals first, check out our Antigravity RAG Pipeline Guide.

Architecture Overview — A Three-Layer Design

The chatbot architecture is organized into three distinct layers, each handling a specific concern.

Presentation Layer (Streaming UI)

This is the frontend layer responsible for user interactions. Using the Vercel AI SDK's useChat hook, it implements Server-Sent Events (SSE) based streaming. As tokens are generated, they're reflected in the UI in real time, giving users an impression of near-instant responses.

Orchestration Layer (Function Calling Router)

This middleware layer mediates between the AI model and external tools. It analyzes user intent, selects the appropriate tool (function), and executes it. By combining multiple tools — weather lookups, database queries, external API calls — you can dramatically extend the AI's capabilities.

Knowledge Layer (RAG Pipeline)

This layer enhances answer accuracy through a knowledge base. Documents are split into chunks, converted to vector embeddings, and stored in a vector database. When a user asks a question, semantically similar documents are retrieved and passed as context to the LLM, significantly reducing hallucinations.

// Conceptual architecture structure
// Presentation Layer → Orchestration Layer → Knowledge Layer
 
interface ChatbotArchitecture {
  // Streaming UI Layer
  presentation: {
    framework: "Next.js App Router";
    streaming: "Vercel AI SDK useChat";
    transport: "Server-Sent Events (SSE)";
  };
  // Function Calling Router
  orchestration: {
    model: "Gemini 2.5 Pro" | "Claude 4 Sonnet";
    tools: ToolDefinition[];
    router: "AI-driven tool selection";
  };
  // RAG Pipeline
  knowledge: {
    embedding: "@cf/baai/bge-m3";  // multilingual, handles Japanese
    vectorDB: "Cloudflare Vectorize" | "Pinecone";
    chunking: "semantic-splitting";
  };
}
✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦Why 75.4% of the corpus was silently discarded in Japanese, and the tokenizer-based chunker that fixed it
✦Embedding model selection on real numbers (512 vs 60,000 token ceiling, 16.7x price gap) and why both sides of the pipeline must match
✦A 7-item pre-production checklist (embedding cost, rate limiting, timeout design) hardened in real operation
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

⚙ AI Tools2026-06-12
Cutting Down 'Plausible but Wrong' RAG Answers — A Retrieval Evaluation Harness for Gemma 4 and Antigravity
Replace gut feeling with recall@5, MRR and faithfulness scores — a 30-question golden dataset and a small Python harness for evaluating a local Gemma 4 RAG stack.
⚙ AI Tools2026-04-21
Prompts Are Assets: Building a Production-Grade Prompt Management Platform with Antigravity — Versioning, A/B Testing, and Quality Evaluation
A hands-on implementation guide for treating prompts as first-class code — with versioning, A/B testing, and automated quality evaluation. Design patterns and working code for running AI agents on Antigravity with safe, continuous prompt improvement.
⚙ AI Tools2026-10-01
You Don't Need to Read Commands: Three Places to Check When an Agent Works
If strings of terminal text make the approve button feel risky, here is a way to check just three things: a way back before you start, the first word of a command, and the diff afterward.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links