How Semantic Caching Works
Picture a mobile app with hundreds of thousands of users. Every day it generates thousands of similar LLM queries: “How to add a contact?”, “How to create a new contact?”, “How to enter a contact in the list?”. Without semantic caching, each such request hits the API, multiplying costs. We solve this by deploying a mechanism that stores responses together with the vector representation of the query. On repeat requests, the system searches for semantically close embeddings and returns the stored answer, bypassing the LLM. In practice, this reduces API costs by 40–60% and cuts latency from seconds to milliseconds.
Problems Semantic Cache Solves
Standard exact-key caching is powerless against synonyms and rephrasings. Users phrase the same question differently — and you pay each time. LLM APIs are expensive at high frequency of repetitive requests; a typical hit rate without caching is near zero. Response latency of 2–5 seconds hurts UX in a mobile app, especially on slow channels. Semantic cache solves all three problems at once.
How We Do It: FastAPI Middleware
The server side is FastAPI middleware. On each request, we generate an embedding via OpenAI, search for the nearest one in a vector store. If cosine similarity exceeds a threshold, we return the cache; otherwise, we call the LLM and save the new embedding. Example code:
import numpy as np from openai import AsyncOpenAI client = AsyncOpenAI() cache: list[dict] = [] # In production: Redis + pgvector or Pinecone async def get_embedding(text: str) -> list[float]: response = await client.embeddings.create( model="text-embedding-3-small", input=text ) return response.data[0].embedding def cosine_similarity(a: list[float], b: list[float]) -> float: a_arr, b_arr = np.array(a), np.array(b) return float(np.dot(a_arr, b_arr) / (np.linalg.norm(a_arr) * np.linalg.norm(b_arr))) async def semantic_cache_lookup(query: str, threshold: float = 0.92) -> str | None: query_emb = await get_embedding(query) for entry in cache: similarity = cosine_similarity(query_emb, entry["embedding"]) if similarity >= threshold: return entry["response"] return None Threshold is a critical parameter. At 0.85 the cache becomes too aggressive: semantically different questions get the same answer. At 0.97 it barely works. The optimal range for most domains is 0.90–0.95, tuned on real queries.
Step-by-Step Semantic Cache Setup
- Log user queries in production (at least 1000).
- Generate embeddings using the chosen model (we use text-embedding-3-small).
- Build a vector index: HNSW for fast search.
- Tune threshold on a holdout set: analyze hit rate and quality.
- Deploy middleware on the server side.
- Monitor hit rate, savings, and false positives.
Why Threshold 0.92 Is an Optimal Start
At threshold 0.92, false positives are minimal, and hit rate on typical questions reaches 40–60%. Lower values give more matches but degrade answer quality. Higher values sharply reduce cache effectiveness. We always fine‑tune the exact value on your logs for the best balance.
When Semantic Cache Does Not Work
For dynamic data — user balance, order status, exchange rates — caching is useless. We identify such requests with a classifier and exclude them from the cache. Also, caching is ineffective if questions are unique and never repeated.
Redis vs pgvector: Which to Choose
Redis with RediSearch is 3× faster than pgvector for caches up to 50k entries, but pgvector scales to millions without precision loss.
| Storage | Performance (latency) | Scaling | Setup Complexity |
|---|---|---|---|
| Redis + RediSearch | 1–5 ms for 50k entries | Medium (up to 100k) | Low |
| pgvector (PostgreSQL) | 5–15 ms for 100k entries | High (millions) | Medium |
| Pinecone (managed) | 2–10 ms | Very high | Low |
| Embedding Model | Dimensions | Price per 1K tokens | Accuracy on our domain |
|---|---|---|---|
| text-embedding-3-small | 1536 | $0.13 | 0.92 |
| text-embedding-3-large | 3072 | $0.25 | 0.97 |
For cosine similarity we use the standard formula: cosine of the angle between vectors via dot product.
Invalidation and TTL
Semantic cache must be invalidated when the system prompt or base model is updated — old answers may not match new behavior. Recommended TTL: 7–30 days for stable FAQ-like questions. For time‑sensitive questions, we do not apply caching.
What’s Included in the Work
- Architecture diagram for integrating semantic cache into a mobile app (iOS/Android).
- Setup of embedding generation and model selection (OpenAI, Cohere, SentenceTransformers).
- Tuning similarity threshold on your query logs.
- Implementation of server‑side middleware (FastAPI, Node.js, Go).
- Monitoring of hit rate and cost savings.
- Documentation and team training.
Our experience includes 5+ deployments for apps with audience from 10k to 1M DAU. We guarantee a hit rate of at least 40% on stable questions.
Typical Mistakes When Implementing
- Choosing too low a threshold — the cache starts confusing semantically different queries.
- Ignoring invalidation when the prompt changes — users get outdated responses.
- Lack of fallback: if vector search fails, the request must go directly to the LLM.
- Incorrect choice of vector index (flat vs HNSW) for the cache size.
Timeline Estimates
Basic semantic cache on Redis + OpenAI Embeddings — 2–3 days. With threshold tuning on real data and hit‑rate monitoring — 3–5 days. If integration into existing mobile infrastructure is needed, contact us for an estimate. Also, order an audit of your current AI costs — we will calculate potential savings and propose an architecture for your load. Get a consultation from an engineer who has already deployed such solutions.
Source: Redis Stack documentation







