Antigravity × Custom AI Chatbot — Notes from Building RAG, Function Calling, and Streaming UI
Build notes for a custom AI chatbot with RAG, Function Calling, and Streaming UI, including the token-estimation bug that silently discarded 75% of the corpus.
Setup and context — the recall problem that started this
When I wired a small RAG chatbot into the support flow for one of my wallpaper apps, the answers were mediocre for the first few weeks. Roughly 1,200 FAQ entries were sitting in the index, yet asking "how do I cancel" would not surface the cancellation steps in the top results. The embedding model and the vector database were both configured straight from the official samples. Nothing threw an error.
It took me an embarrassingly long time to find the cause. It was one line in the chunker: Math.ceil(text.length / 4). That estimate is off by roughly 3.9x for Japanese text. Chunks I believed were capped at 512 tokens were actually running past 2,000 — and the embedding model's input limit is 512. About three quarters of every document I fed in was being discarded silently, with no error and no warning.
What follows is the full build, including that detour: RAG, Function Calling, and Streaming UI combined into one working stack, with the code you need to run it.
RAG (Retrieval-Augmented Generation): Searches your own documents and databases to improve answer accuracy and reduce hallucinations
Function Calling: Dynamically connects to external APIs and databases to fetch real-time information or perform actions
Streaming UI: Displays token-by-token responses in real time, dramatically improving perceived response speed
This article is aimed at engineers with experience building AI applications, assuming familiarity with TypeScript, Next.js, and vector databases. If you'd like to learn RAG fundamentals first, check out our Antigravity RAG Pipeline Guide.
Architecture Overview — A Three-Layer Design
The chatbot architecture is organized into three distinct layers, each handling a specific concern.
Presentation Layer (Streaming UI)
This is the frontend layer responsible for user interactions. Using the Vercel AI SDK's useChat hook, it implements Server-Sent Events (SSE) based streaming. As tokens are generated, they're reflected in the UI in real time, giving users an impression of near-instant responses.
Orchestration Layer (Function Calling Router)
This middleware layer mediates between the AI model and external tools. It analyzes user intent, selects the appropriate tool (function), and executes it. By combining multiple tools — weather lookups, database queries, external API calls — you can dramatically extend the AI's capabilities.
Knowledge Layer (RAG Pipeline)
This layer enhances answer accuracy through a knowledge base. Documents are split into chunks, converted to vector embeddings, and stored in a vector database. When a user asks a question, semantically similar documents are retrieved and passed as context to the LLM, significantly reducing hallucinations.
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦Why 75.4% of the corpus was silently discarded in Japanese, and the tokenizer-based chunker that fixed it
✦Embedding model selection on real numbers (512 vs 60,000 token ceiling, 16.7x price gap) and why both sides of the pipeline must match
✦A 7-item pre-production checklist (embedding cost, rate limiting, timeout design) hardened in real operation
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
The most critical factor in RAG pipeline quality is the document chunking strategy. Rather than simple character-count splitting, we adopt "semantic chunking" that splits at meaningful boundaries.
Let Antigravity's agent generate the following chunking module after explaining your project structure:
// src/lib/rag/chunker.ts// Semantic chunking for document preprocessinginterface DocumentChunk { id: string; content: string; metadata: { source: string; section: string; position: number; tokenCount: number; }; embedding?: number[];}const CHUNK_CONFIG = { maxTokens: 512, // Maximum tokens per chunk overlapTokens: 64, // Overlap between chunks minTokens: 100, // Minimum tokens (merge into previous if below)} as const;export function splitDocumentSemantically( text: string, source: string): DocumentChunk[] { // Primary split at section boundaries (headings) const sections = text.split(/(?=^#{1,3}\s)/m); const chunks: DocumentChunk[] = []; let position = 0; for (const section of sections) { const sectionTitle = section.match(/^#{1,3}\s(.+)/)?.[1] ?? "untitled"; const paragraphs = section.split(/\n\n+/); let currentChunk = ""; for (const para of paragraphs) { const combined = currentChunk ? `${currentChunk}\n\n${para}` : para; const tokenEstimate = Math.ceil(combined.length / 4); if (tokenEstimate > CHUNK_CONFIG.maxTokens && currentChunk) { chunks.push({ id: `${source}-${position}`, content: currentChunk.trim(), metadata: { source, section: sectionTitle, position: position++, tokenCount: Math.ceil(currentChunk.length / 4), }, }); // Overlap: include tail of previous chunk at start of next const overlapText = currentChunk.slice( -(CHUNK_CONFIG.overlapTokens * 4) ); currentChunk = `${overlapText}\n\n${para}`; } else { currentChunk = combined; } } if (currentChunk.trim()) { chunks.push({ id: `${source}-${position}`, content: currentChunk.trim(), metadata: { source, section: sectionTitle, position: position++, tokenCount: Math.ceil(currentChunk.length / 4), }, }); } } return chunks;}
The token-estimation trap — length / 4 is off by ~3.9x in Japanese
The code above still contains the exact landmine I stepped on: Math.ceil(combined.length / 4).
"Roughly four characters per token" is a rule of thumb you see all over English-language material, and for English it holds up fine. For Japanese it collapses. Here is what I measured with the cl100k_base tokenizer.
Language
Characters per token (measured mean)
Error vs. the / 4 assumption
Japanese (5 samples)
1.02
underestimates by ~3.9x
English (5 samples)
5.49
0.73x — errs on the safe side
In Japanese, roughly one character is one token. A chunk you sized with / 4 is carrying close to four times the tokens you think it is. This is precisely why porting an English tutorial verbatim hides the bug: in English the estimate errs high, so nothing ever breaks.
The damage shows up when that miscount meets the model's input ceiling. Cloudflare's official model info for @cf/baai/bge-large-en-v1.5 states Maximum Input Tokens: 512. Anything past that is not an exception — it is silently dropped.
I ran a 2,018-character Japanese document (2,084 tokens, measured) through both splitters.
Splitting logic
Chunks
Tokens per chunk
Chunks over 512
Tokens discarded by the model
estimate via length / 4
1
2,084
1 / 1
1,572 (75.4%)
count with the tokenizer
5
352 / 447 / 443 / 429 / 413
0 / 5
0 (0.0%)
Three quarters of the corpus never reached the index at all. Nothing in the logs said so. No amount of tuning on the retrieval side was ever going to find it.
The fix is simply to stop estimating and start counting.
// src/lib/rag/chunker.ts (corrected)// Count real tokens instead of guessing from character lengthimport { get_encoding } from "tiktoken";// Cut at 448 against a model ceiling of 512, leaving room for the// overlap and for the occasional long proper noun that blows up the count.const CHUNK_CONFIG = { maxTokens: 448, overlapTokens: 64, minTokens: 100, hardLimit: 512, // Maximum Input Tokens for @cf/baai/bge-large-en-v1.5} as const;const enc = get_encoding("cl100k_base");export function countTokens(text: string): number { return enc.encode(text).length;}export function splitDocumentSemantically( text: string, source: string): DocumentChunk[] { const sections = text.split(/(?=^#{1,3}\s)/m); const chunks: DocumentChunk[] = []; let position = 0; for (const section of sections) { const sectionTitle = section.match(/^#{1,3}\s(.+)/)?.[1] ?? "untitled"; const paragraphs = section.split(/\n\n+/); let currentChunk = ""; for (const para of paragraphs) { const combined = currentChunk ? `${currentChunk}\n\n${para}` : para; if (countTokens(combined) > CHUNK_CONFIG.maxTokens && currentChunk) { chunks.push(makeChunk(currentChunk, source, sectionTitle, position++)); // Take the overlap in tokens too — slicing by characters // walks straight back into the same trap. currentChunk = `${takeLastTokens(currentChunk, CHUNK_CONFIG.overlapTokens)}\n\n${para}`; } else { currentChunk = combined; } } if (currentChunk.trim()) { chunks.push(makeChunk(currentChunk, source, sectionTitle, position++)); } } // Last line of defense: never let an oversized chunk through quietly. for (const c of chunks) { if (c.metadata.tokenCount > CHUNK_CONFIG.hardLimit) { throw new Error( `chunk ${c.id} is ${c.metadata.tokenCount} tokens (limit ${CHUNK_CONFIG.hardLimit})` ); } } return chunks;}function makeChunk( content: string, source: string, section: string, position: number): DocumentChunk { const trimmed = content.trim(); return { id: `${source}-${position}`, content: trimmed, metadata: { source, section, position, tokenCount: countTokens(trimmed) }, };}function takeLastTokens(text: string, n: number): string { const ids = enc.encode(text); if (ids.length <= n) return text; return new TextDecoder().decode(enc.decode(ids.slice(-n)));}
Of everything here, the piece I value most is that final throw. Truncation is dangerous precisely because it happens quietly, so the pipeline should be the thing that makes noise. Put that check in once and a future change to your document formatting cannot drop you back into the same hole.
Choosing the embedding model — don't point -en- at Japanese
With the chunker fixed, the model itself was next. I had started on @cf/baai/bge-large-en-v1.5, and the -en- in that model ID is not decoration. Japanese text runs through an English model without complaint and vectors come back, but whether those numbers capture Japanese meaning is a separate question.
Workers AI also offers @cf/baai/bge-m3, described on its official page as a "Multi-Functionality, Multi-Linguality, and Multi-Granularity embeddings model." Lining the two up using Cloudflare's own model info left very little to deliberate over.
Property
@cf/baai/bge-large-en-v1.5
@cf/baai/bge-m3
Language
English-oriented (-en- in the model ID)
Multilingual
Input ceiling
Maximum Input Tokens 512
Context window 60,000 tokens
Unit price
$0.20 per M input tokens
$0.012 per M input tokens
Batch limit
Up to 100 items per request
Batch supported
That is a 16.7x price difference. Re-indexing a two-million-character Japanese knowledge base (about 1.957M tokens, measured) costs roughly $0.39 on the first model and about $0.023 on the second. The absolute numbers are small, but the ratio scales directly with your corpus. Going from a 512-token ceiling to a 60,000-token context window also buys real freedom in how you size chunks.
Cheaper, multilingual, and a far higher ceiling — an unusually easy call. I moved to bge-m3 here.
One caveat: the bge-m3 documentation lists query and contexts[] in its parameter table while the code samples on the same page send { text: [...] }. The docs contradict themselves. When you migrate, push a handful of items through first and look at the shape of what comes back before you run the whole corpus.
Vector Embeddings and Index Building
Once documents are chunked, we convert them into vector embeddings and store them in a vector database. This example uses Cloudflare Vectorize, but the same pattern applies to Pinecone or Weaviate.
// src/lib/rag/embedder.ts// Vector embedding generation and index storageinterface EmbeddingResult { chunkId: string; vector: number[]; dimensions: number;}// A Vectorize index fixes its dimensionality at creation time. If this// doesn't match the model's real output width, every upsert fails.// bge-large-en-v1.5 is 1,024 dimensions per Cloudflare's docs — not 768.const EMBEDDING_DIM = 1024;export async function generateEmbeddings( chunks: DocumentChunk[], env: { AI: Ai }): Promise<EmbeddingResult[]> { // Use Cloudflare Workers AI embedding model // The official schema caps the text array at 100 items; 50 leaves headroom. const batchSize = 50; const results: EmbeddingResult[] = []; for (let i = 0; i < chunks.length; i += batchSize) { const batch = chunks.slice(i, i + batchSize); const texts = batch.map((c) => c.content); // Use the multilingual model. Don't point bge-large-en-v1.5 at Japanese. const response = await env.AI.run( "@cf/baai/bge-m3", { text: texts } ); // Expected output: { data: Array<number[]> } for (let j = 0; j < batch.length; j++) { results.push({ chunkId: batch[j].id, vector: response.data[j], dimensions: EMBEDDING_DIM, }); } } return results;}export async function indexToVectorize( chunks: DocumentChunk[], embeddings: EmbeddingResult[], env: { VECTORIZE: VectorizeIndex }): Promise<void> { const vectors = chunks.map((chunk, i) => ({ id: chunk.id, values: embeddings[i].vector, metadata: { content: chunk.content, source: chunk.metadata.source, section: chunk.metadata.section, }, })); // Vectorize batch limit is 1000 items const insertBatchSize = 1000; for (let i = 0; i < vectors.length; i += insertBatchSize) { await env.VECTORIZE.upsert(vectors.slice(i, i + insertBatchSize)); }}
Query Optimization — Hybrid Search
Simply passing the user's raw question to vector search doesn't always produce accurate results. To improve retrieval quality, we combine query rewriting with hybrid search (vector search + keyword search).
// src/lib/rag/retriever.ts// Hybrid search for context retrievalinterface RetrievalResult { content: string; score: number; source: string; section: string;}export async function retrieveContext( query: string, env: { AI: Ai; VECTORIZE: VectorizeIndex }, options: { topK?: number; scoreThreshold?: number } = {}): Promise<RetrievalResult[]> { const { topK = 5, scoreThreshold = 0.7 } = options; // Step 1: Rewrite query for optimal search performance const rewrittenQuery = await rewriteQuery(query, env); // Step 2: Vector search // Query-side and index-side must use the same model and the same pooling. // Swap one without the other and the vector space diverges — scores stop meaning anything. const queryEmbedding = await env.AI.run( "@cf/baai/bge-m3", { text: [rewrittenQuery] } ); const vectorResults = await env.VECTORIZE.query( queryEmbedding.data[0], { topK: topK * 2, // Fetch extra before score filtering returnMetadata: "all", } ); // Step 3: Filter by score threshold const filtered = vectorResults.matches .filter((m) => m.score >= scoreThreshold) .slice(0, topK) .map((m) => ({ content: m.metadata?.content as string, score: m.score, source: m.metadata?.source as string, section: m.metadata?.section as string, })); return filtered;}async function rewriteQuery( originalQuery: string, env: { AI: Ai }): Promise<string> { // Use LLM to optimize the query for search const response = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { messages: [ { role: "system", content: "Rewrite the user question as a search query optimized for semantic search. Return only the rewritten query, nothing else.", }, { role: "user", content: originalQuery }, ], max_tokens: 200, }); return (response as { response: string }).response || originalQuery;}
Function Calling — Dynamic External Tool Integration
Tool Definitions and Schema Design
Function Calling allows your AI chatbot to dynamically interact with external APIs and databases. The key is writing clear, detailed tool definitions — the AI model reads these to decide which tools to use and when. The schema patterns from our Zod Schema-Driven Development Guide are directly applicable here.
// src/lib/tools/definitions.ts// Tool definitions for AI to useimport { z } from "zod";export const toolDefinitions = { // Product search tool searchProducts: { description: "Search for products in the catalog by name, category, or price range. Use this when the user asks about available products or wants recommendations.", parameters: z.object({ query: z.string().describe("Search query for product name or description"), category: z.string().optional().describe("Product category filter"), minPrice: z.number().optional().describe("Minimum price in USD"), maxPrice: z.number().optional().describe("Maximum price in USD"), limit: z.number().default(5).describe("Maximum number of results"), }), }, // Inventory check tool checkInventory: { description: "Check the current inventory status for a specific product. Use this when the user asks if a product is in stock.", parameters: z.object({ productId: z.string().describe("The product ID to check"), }), }, // Order creation tool createOrder: { description: "Create a new order for the user. Only use this after confirming the product and quantity with the user.", parameters: z.object({ productId: z.string().describe("The product ID to order"), quantity: z.number().min(1).describe("Quantity to order"), shippingAddress: z.string().describe("Delivery address"), }), }, // Knowledge base search (integrates with RAG) searchKnowledgeBase: { description: "Search the internal knowledge base for answers to user questions about policies, procedures, or product documentation.", parameters: z.object({ query: z.string().describe("The user's question to search for"), }), },};
Tool Execution Engine
Beyond defining tools, you need a runtime that actually executes them. Structuring each tool's return value consistently helps the AI interpret results correctly.
Using the Vercel AI SDK (ai package), we implement an SSE-based streaming API. This complete API route integrates both Function Calling and RAG.
// src/app/api/chat/route.ts// Streaming chat API routeimport { streamText } from "ai";import { createGoogleGenerativeAI } from "@ai-sdk/google";import { z } from "zod";const google = createGoogleGenerativeAI({ apiKey: process.env.GOOGLE_AI_API_KEY,});export async function POST(req: Request) { const { messages } = await req.json(); // RAG context retrieval using the latest user message const lastUserMessage = messages .filter((m: { role: string }) => m.role === "user") .pop(); const ragContext = lastUserMessage ? await retrieveContext(lastUserMessage.content, env) : []; // Inject RAG context into system prompt const systemPrompt = buildSystemPrompt(ragContext); const result = streamText({ model: google("gemini-2.5-pro"), system: systemPrompt, messages, tools: { searchProducts: { description: "Search for products in the catalog", parameters: z.object({ query: z.string(), category: z.string().optional(), limit: z.number().default(5), }), execute: async (args) => { const result = await executeTool("searchProducts", args, env); return result.data; }, }, checkInventory: { description: "Check product inventory status", parameters: z.object({ productId: z.string(), }), execute: async (args) => { const result = await executeTool("checkInventory", args, env); return result.data; }, }, }, maxSteps: 5, // Limit on chained tool calls onFinish: async ({ usage }) => { // Log token usage for cost tracking console.log( `Tokens used: ${usage.promptTokens} prompt, ${usage.completionTokens} completion` ); }, }); return result.toDataStreamResponse();}function buildSystemPrompt(ragContext: RetrievalResult[]): string { const basePrompt = `You are an AI assistant that answers questions about products.Provide polite and accurate responses.`; if (ragContext.length === 0) return basePrompt; const contextSection = ragContext .map((r) => `[Source: ${r.source}]\n${r.content}`) .join("\n\n---\n\n"); return `${basePrompt}Below are relevant excerpts from our documentation. Prioritize this information in your answers:${contextSection}Important: If the information isn't found in the documents, honestly say "I wasn't able to confirm that."`;}
Frontend: Real-Time Chat UI
On the frontend, we use the useChat hook to render streaming responses in real time. Intermediate tool call states are also displayed, making the AI's thought process visible to users.
As conversations grow longer, managing the context window (token limit) becomes critical. Sending the entire conversation history with every request causes costs to spike and eventually hits the token limit.
Sliding Window + Summary Strategy
The most effective balance between cost and quality is the "sliding window + summary" approach: keep the most recent N messages verbatim while replacing older messages with a compressed summary.
// src/lib/memory/conversation-manager.ts// Conversation memory — sliding window + summaryinterface Message { role: "user" | "assistant" | "system"; content: string;}interface ManagedConversation { summary: string | null; // Summary of older messages recentMessages: Message[]; // Recent conversation history totalTokens: number;}const MEMORY_CONFIG = { maxRecentMessages: 20, // Number of recent messages to keep maxTokenBudget: 8000, // Token budget for context window summaryTriggerCount: 15, // Message count threshold for summarization} as const;export async function manageConversation( allMessages: Message[], env: { AI: Ai }): Promise<ManagedConversation> { // If message count is below threshold, return everything if (allMessages.length <= MEMORY_CONFIG.maxRecentMessages) { return { summary: null, recentMessages: allMessages, totalTokens: countConversationTokens(allMessages), }; } // Convert older messages into a summary const oldMessages = allMessages.slice( 0, allMessages.length - MEMORY_CONFIG.maxRecentMessages ); const recentMessages = allMessages.slice( allMessages.length - MEMORY_CONFIG.maxRecentMessages ); const summary = await generateSummary(oldMessages, env); return { summary, recentMessages, totalTokens: countConversationTokens(recentMessages) + countConversationTokens([ { role: "system", content: summary }, ]), };}async function generateSummary( messages: Message[], env: { AI: Ai }): Promise<string> { const conversationText = messages .map((m) => `${m.role}: ${m.content}`) .join("\n"); const result = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { messages: [ { role: "system", content: "Summarize the following conversation concisely. Preserve important details like user names, order contents, and question context. Keep it under 200 words.", }, { role: "user", content: conversationText }, ], max_tokens: 300, }); return (result as { response: string }).response;}import { countTokens } from "@/lib/rag/chunker";// Use the same counting path for conversation history. Leave a character// estimate here and you'll size Japanese history at about a quarter of its// real weight — the window never shrinks and the context quietly overflows.function countConversationTokens(messages: Message[]): number { return messages.reduce((sum, m) => sum + countTokens(m.content), 0);}
This is the same trap from the chunking section, hiding in the memory layer. If anything it is harder to spot here: undercounting tokens keeps the sliding window convinced it still has room, so nothing appears wrong until the moment your oldest instructions get pushed out. Counting in one place and calling it from both turned out to be the safe arrangement.
Production Design Patterns
Rate Limiting and Cost Management
In production, abuse prevention and cost management are essential. We implement distributed rate limiting using Cloudflare Workers KV with per-user token budgets.
To handle AI model API outages and timeouts, we implement a multi-layer fallback strategy that automatically switches to backup models when the primary model is unresponsive.
Finally, here's the deployment configuration for your chatbot on Cloudflare Workers. For a broader overview of building AI apps on Cloudflare Workers, see our Antigravity × Cloudflare Workers AI Edge App Guide. The wrangler.toml sets up bindings for Vectorize, D1, and KV.
What the docs don't tell you — lessons from running this in production
Everything above is design and implementation. But once a chatbot runs in production for a while, you hit judgment calls that no documentation covers. Over a few months of embedding a small RAG chatbot into the support flow of my wallpaper and wellness apps, here are the lessons that mattered most.
A chunk size of around 512 tokens turned out to be the practical sweet spot
Official samples often use larger chunks (1,000–1,500 tokens), but for an app dominated by short FAQ-style questions, that actually hurt accuracy. A large chunk packs several topics into one embedding vector, blurring the search target.
Measuring top-3 recall across chunk sizes on my FAQ dataset (~1,200 entries) showed roughly this pattern:
1,200 tokens: ~72% recall
768 tokens: ~81% recall
512 tokens (overlap 64): ~88% recall
256 tokens: ~83% recall (context gets cut, so it drops again)
So "smaller is better" isn't the rule — the sweet spot is the smallest chunk that doesn't sever context. For my use case, 512 tokens with an overlap of 64 was consistently strong. This shifts with your content, so always measure once on your own data.
That recall table had the wrong labels on the x-axis
There is a postscript to the chunk-size comparison above. Rereading it after finding the token-estimation bug was a mildly humbling experience.
The code running those measurements estimated with length / 4. So the setting I labeled "512 tokens" was in fact cutting at 2,048 characters — around 2,000 tokens of Japanese, well past the model's 512 ceiling. Only the first quarter of each chunk ever reached the model.
The recall figures themselves are real; I measured them. What was wrong was the axis. Read in characters, the table looks like this.
Label at the time
Actual split width
What reached the model
Top-3 recall
1,200 tokens
~4,800 characters
first 512 tokens only
~72%
768 tokens
~3,072 characters
first 512 tokens only
~81%
512 tokens
~2,048 characters
first 512 tokens only
~88%
256 tokens
~1,024 characters
essentially all of it
~83%
Laid out this way, half of my original interpretation falls apart. "Smaller isn't automatically better" was a sound conclusion, but the top three rows improved less because the chunks got smaller and more because less of each chunk was being thrown away. Only the 256-token row (about 1,024 characters) delivered nearly its full content, which means that row is the only place the genuine effect — context getting cut, so recall drops — was actually visible.
After switching to tokenizer-based counting I re-ran the same FAQ dataset. The current setting, cutting at 448 tokens, lands at about 91% top-3 recall. Better than the old 88% best, though hardly a transformation. Getting 88% while discarding three quarters of the corpus says more about the material than the method: FAQ content tends to put the answer near the top of the paragraph, so the surviving prefix carried most of the signal. That was luck, not design.
The durable lesson had less to do with optimal chunk size and more with checking that the thing you are measuring is the thing you think you are measuring. Since then I log the token count that actually reached the model before I touch any other RAG parameter.
Multi-model fallback wasn't insurance — it was a daily event
While building, I treated fallback as insurance against rare outages. But aggregating production logs, the share of requests that dropped to fallback from primary timeouts, rate limits, or transient 5xx reached 3–5% during peak hours. Even at tens of thousands of requests per month, that's not negligible.
What helped was routing the fallback to a faster, lighter model. When the high-end primary is congested, everything tends to be congested, so swapping to another same-tier model just stalls again. Escaping immediately to a flash-class lightweight model kept perceived latency far steadier.
Seven things to settle before going to production
Here is a checklist of items I wish I had handled from day one.
Make embedding regeneration explicit (regenerate only on document update, not every time — this changes cost by 10x)
Use two-stage streaming timeouts (time-to-first-token and time-to-completion)
Rate-limit by both user and IP (just one is trivially bypassed)
Always log fallback triggers and review the rate weekly
Manage the conversation memory cap by token count, not message count
Prepare user-facing failure copy (a plain "we're busy right now" reduced bounce)
Break down monthly cost into embedding, inference, and vector search
Item 1 had the biggest impact: switching from regenerating embeddings on every request to regenerating only on update cut embedding API cost by nearly an order of magnitude. Balancing observability and cost maps directly onto the years I've spent weighing ad revenue against operating cost in my app business — get this wrong and a small solo project tips into the red fast.
Summary
If you touch one thing next, make it a single log line rather than the RAG pipeline itself: right before you send a chunk to the embedding model, log the token count as measured by the tokenizer. That alone would have surfaced on day one the problem that cost me several weeks here.
The design rests on four pieces — semantic chunking, Zod-typed Function Calling, conversation memory capped by token count, and fallback that drops to a lighter model. All four sit on top of one assumption: that what reaches the model is what you intended to send. When that foundation is quietly wrong, no amount of careful tuning above it will move your numbers.
Silent truncation, mismatched dimensions, a model swapped on one side of the pipeline but not the other. RAG problems that look like accuracy problems are very often wiring problems. Before you try to make the system smarter, check whether the text is arriving at all.
Thank you for reading all the way through a long one. If it saves someone else the same detour, the time I lost was worth spending.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.