Semantic Caching for LLM Responses in Antigravity — Estimating Your Own Savings Before You Build It
Building a semantic LLM response cache with Antigravity, pgvector, and Gemini. Covers migrating off the retired text-embedding-004, normalizing truncated vectors, and a formula that turns hit rate and unit-cost ratio into an actual savings estimate.
Have you ever stared at your LLM invoice and realized you're paying full price every time a user asks roughly the same question in slightly different words? I ran into exactly this problem when the Gemini bill for a small side-project chatbot quietly tripled over one month. Digging through the logs, the pattern was obvious: hundreds of semantically identical questions a day, each phrased just differently enough to miss my Redis key-value cache.
A traditional exact-match cache is blind to meaning. This guide walks through building a semantic cache — one that uses embeddings to match queries by intent — using Antigravity as your implementation partner, and taking it all the way to a version that survives production. The finished implementation is modest: pgvector, FastAPI, Gemini. What takes longer is the threshold tuning and the pitfalls you want to avoid. I've tried to document the ones I stumbled into personally, so you don't have to.
By the end, you'll have a working implementation you can paste into an existing FastAPI service, a measurement scaffold that makes threshold tuning data-driven rather than gut-feel, and a practical map of the production hazards that separate "works on my laptop" from "pays for itself three times over each month."
Why LLMs need meaning-based caching, not exact matching
Traditional caching treats lookups as string comparisons. HTTP response caches and Redis both key off URLs or query strings, returning the stored value only when the key matches exactly. This works beautifully for classical workloads — serving the same catalog page to a thousand users, caching the result of a database query — because the inputs are identical across requests. LLM users break this assumption by design, because natural language has enormous expressive redundancy.
Consider these four questions from a real support-bot log, all expressing the same intent:
"How do I cancel my subscription?"
"Where can I unsubscribe?"
"I want to stop billing."
"Can you help me quit?"
Four strings, zero overlap. An exact-match cache hits 0% of the time, even though a human support agent would reuse the same answer verbatim for all four. Run them through a text-embedding model, however, and their pairwise cosine similarities fall between 0.89 and 0.94. The premise of semantic caching is that "if two queries sit close in embedding space, they can share a response." Under that premise, three of those four queries can be served from cache instead of a fresh LLM call.
Three concrete benefits come from this. First, obvious cost reduction — you stop paying for repeat inference. For a bot handling 10,000 queries a day at roughly $0.001 per Gemini Flash call, an 80% hit rate saves around $240 per month, which is often enough to fund the Postgres instance and still leave change. Second, latency: LLM calls routinely take 800 to 2,000 milliseconds, while a vector similarity lookup plus a cache read lands under 50. That's not just "faster" — it's the difference between "feels like a chat" and "feels like a form submission." Third, and often overlooked, is response stability. LLMs have bad days, regional outages, and quality drift between model versions. Cache hits reproduce past good answers instead of gambling on fresh ones, which is a quiet but real improvement in consistency.
The trade-off, of course, is that "close enough in meaning" can shade into "different intent." A query asking "how do I cancel my account?" sits in the same neighborhood as "how do I cancel my last order?" in embedding space, and mis-routing between the two is a real product bug. Much of this article is about making that trade-off quantitative rather than hopeful.
One more framing matters before we go deeper. Semantic caching is not a silver bullet — it's best suited to workloads where the same intent is expressed many times with variation. Support bots, FAQ systems, onboarding assistants, and technical documentation search all fit well. Code generation, creative writing, and highly personalized chat do not. If your users rarely repeat intents, spend your engineering budget elsewhere.
Architecture — start minimal, grow deliberately
Before reaching for complexity, nail down the smallest viable version. A semantic cache boils down to five steps:
Receive a user query.
Convert it to an embedding vector.
Run top-1 similarity search in a vector store.
If similarity exceeds the threshold, return the cached response.
Otherwise, call the LLM and store the query-response pair with its embedding.
For the store, I recommend PostgreSQL with the pgvector extension. The reasoning is prosaic: you almost certainly already have a Postgres instance, and pgvector integrates cleanly with the app database you're already backing up, authenticating against, and monitoring. The pgvector RAG pipeline guide covers the basics, but the relevant point here is that HNSW indexing keeps nearest-neighbor lookup under 50ms at millions of rows. Managed alternatives like Pinecone or Upstash Vector are good, but adding another billing line, another API, and another backup process is a real cost. For solo developers and small SaaS teams, pgvector first is the pragmatic default. You can always migrate later if your cache grows past tens of millions of rows, but in practice most semantic caches stabilize around a few hundred thousand entries — far inside pgvector's sweet spot.
For the embedding model, I use Google's gemini-embedding-001. If you've seen text-embedding-004 in older tutorials — and it was the default in nearly all of them — note that it was retired on January 14, 2026. The same code today returns a 404, which took me longer to diagnose than I'd like to admit. Swapping the model name looks like a one-line change, and it isn't: dimensions, normalization, and task types all shift underneath you. The next section clears those three out of the way before we write anything else.
Alternative models like OpenAI's text-embedding-3-small or locally-hosted Gemma variants work too, but the criterion here is throughput and price, not benchmark recall. You don't need state-of-the-art quality to tell "cancel my subscription" apart from "show my invoice"; you need predictable latency and a billing curve that doesn't punish you for caching.
A design choice that matters more than people realize is whether to embed queries synchronously or asynchronously. Synchronous embedding (compute before storing) is simpler and avoids race conditions but adds to write latency. Asynchronous embedding (store the query and response first, embed in a background job) saves a few hundred milliseconds on write but creates a window where a query exists in the cache but can't be matched, which undermines hit rate on bursty workloads. For most production systems, synchronous wins. Revisit this only if your embedding API becomes a bottleneck.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦You can estimate your own savings before writing any code, by decomposing the reduction rate into hit rate and unit-cost ratio
✦You'll get a production-grade pgvector + FastAPI + Gemini implementation that Antigravity can maintain without losing context
✦You'll migrate off the retired text-embedding-004 without getting caught by dimensions, normalization, or task types
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Migrating off text-embedding-004 — three things that caught me
All three of these fail quietly. Nothing raises; retrieval quality or ranking just gets a little worse, and you find out weeks later.
Dimensions go from a fixed 768 to a default 3072.text-embedding-004 returned 768 dimensions. gemini-embedding-001 defaults to 3072. Feeding that into an existing vector(768) column fails loudly, so you'll catch it. The harder question is whether to keep 3072. Cache tables grow as hit rate improves, and HNSW indexes only perform when the vectors fit in memory. Cache matching doesn't need the resolution that RAG retrieval does, so I pass output_dimensionality=768 and truncate. Keeping the schema unchanged is a nice side benefit.
Truncated vectors are not normalized. This is the quiet one. Matryoshka representation learning means the first 768 dimensions still carry the meaning — but the moment you slice, the norm drifts away from 1. Truncating 5,000 synthetic 3072-dimensional unit vectors to their first 768 dimensions, I measured norms spread across 0.4619 to 0.5376 (mean 0.5001), a 1.164x gap between the largest and smallest. Running 1,000 nearest-neighbor lookups over those vectors, the top-1 result chosen by cosine and the top-1 chosen by inner product disagreed 126 times — 12.6%. After L2 re-normalization, they disagreed zero times. These are synthetic vectors, so don't carry the percentage over to real embedding distributions; as grounds for the rule "if you truncate, normalize," it was enough for me.
As long as you use pgvector's cosine operator <=>, the difference never shows up in rankings — the norm divides out. The risk appears when you switch to the inner-product operator <#> for speed. One character changes, and your threshold quietly means something else. Normalizing on the client side removes the trap entirely.
Don't pin the task type on the instance. In RAG, the convention is RETRIEVAL_DOCUMENT when indexing and RETRIEVAL_QUERY when searching. But a semantic cache isn't comparing documents to queries — it compares past queries to the current one, the same kind of text on both sides. Carry the RAG convention over and you embed the same sentence under two different task types, so vectors that should match don't. For caching, use SEMANTIC_SIMILARITY on both the store and lookup paths. If you go through a wrapper library, avoid setting the task type on the client instance: it's easy to think you've overridden it per call when you haven't. Pass it explicitly every time.
Building the minimum viable version in 30 minutes with Antigravity
Let me walk through the actual implementation. I drive Antigravity in Manager mode, ask it to scaffold the skeleton, and then refine specifics by hand. The key to getting useful output is to pin the runtime assumptions up front: Python 3.11, FastAPI, the google-genai SDK, Postgres 16 with pgvector. Without these constraints, Antigravity tends to generate overly abstract code — adapter patterns, pluggable embeddings, "future-proof" interfaces — that looks impressive in review but adds friction for every real change.
Below is the SemanticCache class, scoped to be drop-in for a FastAPI support-bot. It differs from my first version in three places — vector normalization, a per-connection type codec, and a generation fingerprint — each of which is painful to bolt on later, because adding it means rebuilding the cache you already have.
# semantic_cache.py — drop in to an existing FastAPI app# Prereqs: PostgreSQL 16 + pgvector, google-genai SDK, asyncpg, pgvector[asyncpg]# Expected: first call hits the LLM; subsequent similar calls return cached answer at similarity >= 0.92import asyncioimport hashlibimport jsonimport mathimport osimport asyncpgfrom google import genaifrom google.genai import typesfrom pgvector.asyncpg import register_vectorEMBEDDING_MODEL = "gemini-embedding-001"EMBEDDING_DIM = 768 # default is 3072; truncating means normalizing yourselfGENERATION_MODEL = "gemini-2.5-flash"SIMILARITY_THRESHOLD = 0.92 # tune inside 0.90-0.95 based on real dataCACHE_TTL_SECONDS = 60 * 60 * 24 * 7 # invalidate after 1 weekclient = genai.Client(api_key=os.environ["GEMINI_API_KEY"])def l2_normalize(values: list[float]) -> list[float]: """Truncated MRL vectors drift off the unit sphere, so normalize explicitly.""" norm = math.sqrt(sum(v * v for v in values)) if norm == 0.0: raise ValueError("embedding API returned a zero vector") return [v / norm for v in values]def prompt_fingerprint(system_instruction: str, model: str, temperature: float) -> str: """Fold system prompt, model, and temperature into 16 hex chars. Change one, get a new generation.""" raw = json.dumps( {"system": system_instruction, "model": model, "temperature": temperature}, sort_keys=True, ensure_ascii=False, ) return hashlib.sha256(raw.encode("utf-8")).hexdigest()[:16]async def create_pool(dsn: str) -> asyncpg.Pool: """Register the vector codec per connection. Skip this and asyncpg cannot encode a list.""" return await asyncpg.create_pool(dsn, init=register_vector)class SemanticCache: def __init__(self, pool: asyncpg.Pool, fingerprint: str): self.pool = pool self.fingerprint = fingerprint async def _embed(self, text: str) -> list[float]: """SEMANTIC_SIMILARITY on both paths — we compare queries to queries, not to documents.""" result = await asyncio.to_thread( client.models.embed_content, model=EMBEDDING_MODEL, contents=text, config=types.EmbedContentConfig( task_type="SEMANTIC_SIMILARITY", output_dimensionality=EMBEDDING_DIM, ), ) return l2_normalize(result.embeddings[0].values) async def lookup(self, query: str, lang: str) -> tuple[str, float] | None: """Search for a semantically similar cached response. Returns (response, similarity) on hit.""" vec = await self._embed(query) async with self.pool.acquire() as conn: row = await conn.fetchrow( """ SELECT response, 1 - (embedding <=> $1::vector) AS similarity FROM llm_cache WHERE fingerprint = $2 AND lang = $3 AND created_at > NOW() - ($4::int * INTERVAL '1 second') ORDER BY embedding <=> $1::vector LIMIT 1 """, vec, self.fingerprint, lang, CACHE_TTL_SECONDS, ) if row and row["similarity"] >= SIMILARITY_THRESHOLD: return (row["response"], row["similarity"]) return None async def store(self, query: str, lang: str, response: str) -> None: """Persist query/response with its embedding, scoped by generation and language.""" vec = await self._embed(query) async with self.pool.acquire() as conn: await conn.execute( """ INSERT INTO llm_cache (query, lang, fingerprint, embedding, response) VALUES ($1, $2, $3, $4::vector, $5) """, query, lang, self.fingerprint, vec, response, ) async def ask(self, query: str, lang: str = "en") -> dict: """Public entry point. Records both hits and misses.""" hit = await self.lookup(query, lang) if hit: return {"response": hit[0], "cached": True, "similarity": hit[1]} gen = await asyncio.to_thread( client.models.generate_content, model=GENERATION_MODEL, contents=query, ) text = gen.text await self.store(query, lang, text) return {"response": text, "cached": False, "similarity": None}
The matching schema looks like this. lang and fingerprint are columns rather than conventions, because the two pitfalls they guard against — cross-language hits and stale prompt generations — are much easier to stop in the schema than in a code review checklist.
-- schema.sql — run once via psqlCREATE EXTENSION IF NOT EXISTS vector;CREATE TABLE llm_cache ( id BIGSERIAL PRIMARY KEY, query TEXT NOT NULL, lang TEXT NOT NULL, -- structurally prevents cross-language false hits fingerprint TEXT NOT NULL, -- generation of system prompt + model + temperature embedding vector(768) NOT NULL, -- normalized, truncated via output_dimensionality=768 response TEXT NOT NULL, created_at TIMESTAMPTZ DEFAULT NOW(), hit_count INT DEFAULT 0);-- HNSW index for cosine distance (pgvector 0.5+)CREATE INDEX llm_cache_embedding_idx ON llm_cache USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);-- Let generation and language narrow the candidate set firstCREATE INDEX llm_cache_scope_idx ON llm_cache (fingerprint, lang, created_at DESC);
A note on index parameters: m = 16 and ef_construction = 64 are sensible defaults for cache-sized corpora (up to a few million rows). The m parameter is the number of connections per node in the HNSW graph; higher values improve recall but slow builds. The ef_construction parameter affects quality at build time. Once your cache is populated, consider tuning ef_search at query time (via SET LOCAL hnsw.ef_search = 40) — it's the quality/latency knob you actually want to reach for in production, not the index-time parameters.
That's the minimum viable version. Send two semantically similar curl requests and the second one returns cached: true with a similarity somewhere in the 0.93–0.97 range.
Tuning the similarity threshold — balancing hit rate and false positives
This is where most production deployments stand or fall. Lowering the threshold raises hit rate but also raises the risk of serving an answer intended for a slightly different intent — a false hit. Raising it is safer but eats into savings.
My rule of thumb is to start at 0.92 and let real data push it somewhere between 0.90 and 0.95. "Rule of thumb" alone is not a production methodology, though. What actually matters is instrumenting from day one so you can compute the right threshold later. Teams that ship semantic caching without logging usually end up re-shipping it three months later when they realize they can't defend the threshold choice in a postmortem.
The following logger records every lookup decision, including the top similarity score, so you can mine it afterwards.
# similarity_logger.py — log every cache decision for later analysisasync def log_decision( pool: asyncpg.Pool, query: str, top_similarity: float | None, used_cache: bool, response_preview: str,) -> None: """Record every threshold-gated decision so we can back out the right cut-off.""" async with pool.acquire() as conn: await conn.execute( """ INSERT INTO cache_decisions (query, top_similarity, used_cache, response_preview, created_at) VALUES ($1, $2, $3, $4, NOW()) """, query, top_similarity, used_cache, response_preview[:200], )# Run after a week of traffic (directly in psql):# -- Histogram of top-similarity bands, split by cache hit status# SELECT# width_bucket(top_similarity, 0.80, 1.00, 20) AS bucket,# COUNT(*) AS total,# COUNT(*) FILTER (WHERE used_cache) AS cache_hits# FROM cache_decisions# WHERE created_at > NOW() - INTERVAL '7 days'# AND top_similarity IS NOT NULL# GROUP BY bucket ORDER BY bucket;
What you typically see in the 0.88–0.92 band is a cluster of near-misses — queries that share intent but didn't quite clear the bar. Sample them randomly, review a few dozen manually, and pick the lowest threshold where your false-hit rate stays under 1%. Avoid handing this step to the agent entirely. The judgment of whether a cached response is "good enough" for a mis-matched query is product design, not tooling, and it needs a human signoff.
One more heuristic I've found useful: set the threshold 0.01 tighter than your data suggests. The 1% buffer absorbs noise from embedding model updates (they do happen) and from distribution shift as your user base changes. A cache tuned to the knife-edge will surface false hits the first time your product evolves.
Invalidation and privacy — trade-offs to settle at design time
Semantic caches have invalidation problems that don't exist for plain KV caches. They also capture user utterances, which introduces privacy concerns the moment you turn the system on.
Three invalidation strategies are worth knowing. Time-based TTLs — shown above as one week — work well for stable documentation but poorly for volatile data like pricing or inventory. Version-based invalidation adds a cache_version column that flushes the whole cache whenever you change the prompt template, system instructions, or model. Tag-based invalidation segments the cache by tenant or category so you can purge a slice without losing the rest. Most real deployments use a combination: short TTL for volatile topics, version-gated everything, tenant tags for isolation.
On the privacy side, assume user input may contain PII — email addresses, phone numbers, internal identifiers, account numbers — and design for it. A lightweight regex pre-filter that skips PII-bearing queries is the minimum bar. If you serve multiple tenants, go further and scope every cache entry to a tenant ID so tenant A's query can't hit tenant B's cached response. The Antigravity cost-optimization guide touches on this briefly, but the general rule is that cost savings and safety always trade against each other — you pay for isolation in hit rate.
A subtler privacy question is whether to store the raw query text at all. My default is yes, because debugging is impossible without it. But if your compliance requirements forbid persisting user input, you can store only the embedding and a hash of the query, losing debuggability but preserving the functional cache. Whichever way you go, make it a deliberate decision and write it down — someone on your future team will thank you.
# privacy_filter.py — skip PII-bearing queries before they land in the cacheimport rePII_PATTERNS = [ re.compile(r"[\w\.-]+@[\w\.-]+\.\w+"), # email re.compile(r"\b\d{3}-\d{4}-\d{4}\b"), # phone-like re.compile(r"\b\d{4}[\s-]?\d{4}[\s-]?\d{4}[\s-]?\d{4}\b"), # card number-like]def contains_pii(text: str) -> bool: return any(p.search(text) for p in PII_PATTERNS)# Usage before storingif contains_pii(query): # Still call the LLM, but don't persist the query/response pair passelse: await cache.store(query, response)
For a broader take on observability in LLM systems — audit logs, anomaly detection, the full monitoring picture — LLMOps production monitoring is the companion read I'd point you to.
Pitfalls — four production failures I actually caused
This section is where I admit to the mistakes that cost me sleep. You should be able to skip each of these on my behalf.
Pitfall 1: Not including prompt template or system instructions in the cache key.
One day I updated the bot's tone from "formal" to "friendly." New answers used the new tone, but the cache kept handing out the old ones. The lookup was keyed purely on the query embedding, which had no knowledge of the system prompt. The fix is a cache_version column that stores a SHA-256 hash of the system prompt, template, model name, and temperature — and the lookup matches only against rows with the current version. Change any of these, cache entries naturally become invisible until rebuilt. Worth noting: don't store the full prompt, just the hash. Prompts can contain customer-specific context, and replicating them across the cache table becomes a compliance liability.
Pitfall 2: Caching time-sensitive responses.
"What's today's exchange rate?" caches terribly. Returning yesterday's answer today damages trust fast. Two ways to handle it: either use a lightweight classifier (regex or a small LLM call) to detect temporal queries and skip the cache for them, or aggressively shorten the TTL for any query that mentions "today," "latest," "current," and similar markers. I went with the classifier route — it has fewer edge cases and handles phrasing variety better than regex. A small Gemini Flash call costs less than a cent per thousand classifications and catches "at the moment," "right now," "as we speak," and all the other temporal phrasings regex never quite nails.
Pitfall 3: Forgetting embedding cost inverts the savings for tiny queries.
This one surprised me. For short queries, the embedding call can cost more than the LLM call it was meant to save. You pay for the embedding on hits and misses alike, and the cheaper your generation model, the smaller the gap you're arbitraging. On Gemini Flash it's entirely possible for the per-lookup embedding charge to exceed the per-hit savings.
The fix is a two-layer cache: check an exact-match hash first (free), fall through to the semantic layer only on a miss. The next section covers that structure along with what it does to the savings formula.
Pitfall 4: Cross-lingual false hits in multilingual apps.
Multilingual embedding models place semantically equivalent sentences close together regardless of language. This is a feature for RAG, a bug for caching. An English query "How do I cancel?" can hit a Japanese cache entry and return a Japanese response to an English user. The fix is dumb but effective: include a detected language code in the cache key. The langdetect library does the job in a few milliseconds per query. Bonus: the same tag lets you serve different cached responses per language when your underlying knowledge base is regional — legal disclaimers and pricing often differ by locale in ways the LLM cannot know.
Testing and evaluation — building the quality safety net
One question that comes up as soon as you ship this to staging: how do you verify the cache isn't degrading quality? An untested semantic cache is essentially a random-response generator with a 1% failure rate, and that failure rate compounds in scary ways across user sessions.
The minimum viable evaluation approach is a golden dataset. Collect 100 to 300 real queries from your logs, hand-annotate the correct answer or accepted answer range, and re-run the dataset against your cache periodically. Compare the cached answer with the annotated one using either exact-match, a bleu-like overlap score, or — most accurately — an LLM-as-judge prompt that grades whether the cached answer is substantively equivalent to the annotation. Run this nightly and alert on regressions.
A more ambitious approach is shadow testing. Route every cache hit through a background thread that also fires the real LLM call, then compare the cached and fresh responses. This doubles your LLM spend during the evaluation window, so keep it to a sampled percentage (say, 1% of traffic). The payoff is a continuous signal about how often cached responses diverge from what the fresh LLM would say, which is the closest you can get to "is this cache still safe" without asking users.
For both approaches, the judgment layer matters. Exact-match is too strict because LLM output has natural variation; a character-by-character comparison will flag tolerable paraphrases as mismatches. Semantic comparison via another embedding call is closer but still coarse. The honest answer is that an LLM judge with a carefully written rubric works best, at the cost of another tokens-per-evaluation line item. Budget for this, because the alternative is learning about quality regressions from customer complaints.
Scaling considerations — from one instance to one thousand
The implementation above works for a single Postgres instance serving a single application. A few scaling notes for when you outgrow that.
First, the cache table will grow. At 10,000 queries per day and a 1-week TTL, you're looking at 70,000 rows plus embeddings — roughly 250MB of storage. That's fine. At 1 million queries per day, you're looking at 7 million rows and around 25GB, which strains commodity HNSW indexes. Two mitigations: shorten the TTL (if most hits come in the first day, a 24-hour TTL may lose almost no value) and consider partitioning by tenant or by time window, pruning old partitions rather than doing per-row deletes.
Second, read-heavy cache workloads benefit from read replicas. Point the lookup call at a read replica, keep store on the primary, and you can scale horizontally without much fuss. Watch for replication lag — a write followed immediately by a read on the same query can miss the cache until the replica catches up. In practice this is rare enough (replicas typically lag under a second) that it's not worth engineering around.
Third, hot keys still matter. If 5% of your queries produce 60% of the hits, the lookup for those queries is your real bottleneck, not the index. Warm a small in-memory LRU cache in front of the pgvector lookup for the top hundred queries, refreshed every minute. This is boring, but it gets you the last 10x of latency improvement.
Finally, for genuinely large deployments — tens of millions of cache entries — consider migrating from pgvector to a purpose-built vector DB like Qdrant or Weaviate. The break-even in my experience is somewhere around 50 million vectors. Below that, pgvector's operational simplicity usually wins.
Estimating your own savings — two layers and the unit-cost ratio
"How much will this actually save?" is the question you want answered before you build, not after. I spent a while reading other people's numbers and found them close to useless, because the conditions behind them were never stated. Written as a formula, you can just plug in your own.
Start with the two-layer structure. The exact-match layer is a hash lookup, so it costs nothing in embedding charges.
# two_layer.py — put an exact-match layer in front of the semantic oneimport hashlibdef exact_key(query: str, fingerprint: str, lang: str) -> str: """Normalize whitespace first, or trivial spacing differences fragment the cache.""" normalized = " ".join(query.split()) payload = "\x1f".join([fingerprint, lang, normalized]) return hashlib.sha256(payload.encode("utf-8")).hexdigest()async def ask_two_layer(cache, redis, query: str, lang: str = "en") -> dict: key = exact_key(query, cache.fingerprint, lang) cached = await redis.get(key) if cached is not None: return {"response": cached, "cached": True, "layer": "exact"} result = await cache.ask(query, lang) # embedding cost starts here await redis.set(key, result["response"], ex=60 * 60 * 24) result["layer"] = "semantic" return result
That " ".join(query.split()) matters more than it looks. Running "cancel my plan " and "cancel my plan" through it lands on the same key, which is exactly the kind of near-duplicate that would otherwise sail past a naive hash and burn an embedding call.
With that in place, the reduction rate decomposes cleanly. Let h be the semantic hit rate, e the share of queries the exact layer absorbs, and c the cost of one embedding relative to one LLM call:
reduction = h − (1 − e) × c
The first term is what you save; the second is the embedding you pay for on every lookup, hit or miss. You can read c straight off your own invoice once you know your average token count. A few combinations:
Hit rate h
Exact layer e
Cost ratio c
Reduction
0.60
0
0.005
59.5%
0.70
0
0.05
65.0%
0.82
0
0.005
81.5%
0.82
0
0.20
62.0%
0.82
0.75
0.20
77.0%
0.90
0
0.05
85.0%
Two things fall out of this. While c stays small, the reduction is essentially just your hit rate. Push the generation side onto a cheap model until c reaches 0.20, and the same hit rate loses close to 20 points. Compare rows four and five: adding the exact-match layer pulls 62.0% back up to 77.0%. The cheaper your generation model, the more the second layer is worth.
So when you see "we cut costs 80%," read it as a statement about a specific c and a specific h. Mine happened to meet those conditions. Checking the conditions first is still the right order.
Monitoring and A/B testing — turning savings into numbers
Shipping isn't done until you can explain savings as a number to someone else. Four metrics I instrument on day one of every deployment:
Cache hit rate, broken down by time of day and tenant.
Similarity-score histogram so you can see whether your threshold sits in the right place.
Inferred cost savings, computed as hit count times average token cost avoided.
Response latency p50 and p95, split by cache hits versus misses.
Exporting these to Prometheus is a good place to let Antigravity do the boilerplate. Ask it to "add Prometheus metrics to this Python code" and it produces reasonable scaffolding. What you must specify yourself is what to measure. Left to its own devices, the agent tends to standardize on RED (Rate / Error / Duration) and drop domain-specific signals like hit rate and similarity distribution. This is the thread you'll find repeatedly: agents are excellent at boilerplate, weak at deciding which boilerplate is relevant.
For A/B testing, route users to cache-on versus cache-off groups by hashing their user ID, then run both for a week and compare user satisfaction (thumbs-up rate), re-ask rate (how often users re-ask within a minute), average response time, and daily cost. In my case satisfaction stayed within ±1 point and p50 latency dropped from 1.4 seconds to under 300ms. Cost behaved exactly as the formula in the previous section predicts: the weeks where savings landed near 80% were the weeks where hit rate sat in the low eighties and the unit-cost ratio was small. That number describes how repetitive my traffic was that week, not how good the cache is — which is the reason to run the formula on your own numbers rather than borrowing anyone's headline figure, mine included.
One caution on A/B testing: a week may not be enough if your traffic has weekly seasonality (enterprise bots are notorious for this — Monday mornings look nothing like Friday afternoons). Run the test for at least two full weeks if you can afford the wait, or run it continuously with a small sample (say, 5% cache-off) as a permanent canary.
One concrete next step
That's a lot of ground. Here's the single thing worth doing today. Sample 500 to 1,000 recent user queries from your current bot or support system, and count duplicates by eye. If more than 30% feel like "I've seen this question before," the investment in a semantic cache pays back almost certainly. If less than 10%, there's probably a higher-ROI target elsewhere — prompt compression, model downgrading, or batching.
Following the structure above, the minimum viable version takes an afternoon, and monitoring plus threshold tuning fits into a two- or three-day sprint. The real insight isn't that semantic caching is clever — it's that building it alongside the measurement scaffolding, not after, is what lets you iterate fast. Numbers for hit rate, similarity distribution, and false-hit rate, visible in a dashboard, turn the whole thing from a guess into a system you can actually improve.
Share
Thank You for Reading
Antigravity Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.