The AI Caching Playbook, Part 1: The 10 Core Caches
KV, prompt, semantic, embedding, retrieval, reranking, tool, API, conversation and agent state caches: what each saves and how to key it safely.

On this page
Modern AI applications are rarely limited by model intelligence alone.
In production, the real problems are often:
- Why is this request taking 8 seconds?
- Why did our LLM bill double?
- Why are we computing the same embedding thousands of times?
- Why is the agent calling the same API repeatedly?
- Why are users asking the same question, but every request still hits the model?
- Why does our RAG pipeline perform the same retrieval work again and again?
- Why does a long conversation become increasingly expensive?
The answer to many of these problems is caching.
Caching in AI is not one single technology. There are multiple cache layers operating at different points in the AI stack. A production RAG or agentic AI system may look like this:
- User
- Request cache
- Semantic cache
- Prompt cache
- LLM inference
- RetrievalTools
- Retrieval cacheTool result cache
- Reranker cacheAPI response cache
- Final answer
The important idea is:
Don't cache only the final answer. Cache expensive intermediate work too.
This article covers the first 10 cache layers every AI engineer should understand, in the order they sit in the stack:
- KV cache
- Prompt cache
- Semantic cache
- Embedding cache
- Retrieval cache
- Reranking cache
- Tool result cache
- API response cache
- Conversation cache
- Agent state cache
1. KV cache
1.1 What is it?
KV cache stands for key-value cache. It is one of the most important caches inside transformer inference.
When an LLM processes tokens, the attention mechanism creates keys and values for each one. For tokens that have already been processed, those keys and values don't need to be recomputed from scratch during autoregressive generation. The model can reuse them. Conceptually:
- Tokens 1…n
- Attention
- KV cache
- Next token
Without KV caching, every new token means recomputing the whole previous context:
- Generate token 1 → recompute previous context
- Generate token 2 → recompute previous context
- Generate token 3 → recompute previous context
- …and so on
With KV caching, the previous context is processed once and reused:
- Process previous context once
- KV cache
- Reuse during generation
This dramatically improves autoregressive decoding efficiency.
1.2 Where does the KV cache live?
Typically in GPU memory, next to the model weights and the activations. That makes the KV cache fundamentally an inference-engine and model-serving concern.
You generally don't implement KV caching yourself when using a managed LLM API. Inference engines such as vLLM, TensorRT-LLM and other serving systems manage it.
1.3 Why should RAG engineers care?
Because RAG increases context. Every request carries:
- User query
- System prompt
- Retrieved documents
- Conversation history
The larger the context, the more important efficient attention and KV management become. For long-context applications running at 10K, 50K or 100K+ tokens, KV memory can become a major infrastructure constraint.
1.4 Production considerations
Watch:
- KV cache memory usage
- Context length
- Concurrent requests
- GPU memory
- Batch size
- Prefix reuse
- Cache eviction
- Quantization
- Paged KV cache
A production system should treat KV memory as a capacity-planning problem.
2. Prompt caching
2.1 What is it?
Prompt caching is different from KV caching. The idea is:
If a large portion of a prompt stays unchanged, don't repeatedly process the same prefix.
Consider an agent whose prompt contains system instructions, company policies, tool definitions, a large knowledge context and, at the end, the user's question. Only the user's question changes. Every request could look like:
- 100K tokens of static context
- 20 tokens of new user input
Processing that static prefix again and again is wasteful. Prompt caching lets the provider or inference system reuse the previously processed prefix, where supported.
2.2 Example
| Request | Without caching | With prompt caching |
|---|---|---|
| 1 | Process 100K static tokens + 20 new tokens | Process 100K static tokens (cached) + 20 new tokens |
| 2 | Process 100K static tokens again + 20 new tokens | Reuse cached prefix + 20 new tokens |
| 3 | Process 100K static tokens again + 20 new tokens | Reuse cached prefix + 20 new tokens |
2.3 Ideal candidates
Prompt caching works particularly well for:
- Large system prompts. "You are an enterprise support agent…" followed by a long policy document and detailed instructions.
- Agent tool definitions. Tool 1, tool 2, tool 3… tool 50: the same schemas on every call.
- Stable enterprise context. Company policies, product documentation, security rules, compliance instructions.
2.4 Production rule
Put stable content before dynamic content when your model or provider's caching benefits from prefix reuse.
- System prompt
- Company policy
- Tool definitions
- Knowledge
- User query
- User query
- System prompt
- Random context
- Tool definitions
3. Semantic cache
3.1 What is it?
This is one of the most useful caches for RAG applications.
Traditional caching asks:
"Is the request exactly the same?"
Semantic caching asks:
"Is this request meaningfully similar to something we've already answered?"
Consider three questions:
- "What is the refund policy?"
- "Can I get my money back?"
- "How does your refund process work?"
They aren't identical strings, but they may have the same intent. A normal cache treats "What is the refund policy?" and "Can I get my money back?" as two different keys. A semantic cache embeds the question, runs a similarity search and finds the existing answer:
- "What is the refund policy?"
- Embedding
- Similarity search
- Existing answer
3.2 Basic architecture
- User query
- Generate embedding
- Search cache
- Similar enoughreturn cached responseNo matchcontinue
- RAG / LLM
- Store response
3.3 Python example
A simplified implementation can use an embedding model and a list (in production, a vector database):
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("all-MiniLM-L6-v2")
cache = []
def similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
def semantic_cache_lookup(query, threshold=0.90):
query_embedding = model.encode(query)
for item in cache:
score = similarity(query_embedding, item["embedding"])
if score >= threshold:
return item["response"]
return NoneStore a response:
def store_cache(query, response):
embedding = model.encode(query)
cache.append({
"query": query,
"embedding": embedding,
"response": response,
})Then:
query = "Can I get a refund?"
cached = semantic_cache_lookup(query)
if cached:
answer = cached
else:
answer = call_llm(query)
store_cache(query, answer)3.4 The dangerous part
Semantic cache thresholds are extremely important.
- Too low (
threshold = 0.70): you may return incorrect answers. - Too high (
threshold = 0.99): you barely get cache hits.
Production systems should tune thresholds using real traffic.
4. Embedding cache
4.1 What is it?
Embedding generation is another place where repeated computation happens.
Suppose you ingest document.pdf and split it into 10,000 chunks, then generate an embedding for each chunk. If the same chunk is processed again, generating its embedding again is unnecessary.
4.2 Architecture
- Text
- Hashtext + embedding model + version
- Embedding cache
- Hitreturn embeddingMisscall embedding model
- Store in cache
4.3 Simple Python implementation
import hashlib
embedding_cache = {}
def cache_key(text, model_name):
value = f"{model_name}:{text}"
return hashlib.sha256(value.encode()).hexdigest()
def get_embedding(text, model_name, embed_fn):
key = cache_key(text, model_name)
if key in embedding_cache:
return embedding_cache[key]
embedding = embed_fn(text)
embedding_cache[key] = embedding
return embedding4.4 Important production detail
The cache key should not be hash(text) alone. It should usually incorporate:
- the text
- the embedding model
- the embedding model version
- the preprocessing version
Otherwise:
- Old embedding model
- Old cached vector
- New embedding model
- Inconsistent vector space
That can silently damage retrieval quality.
5. Retrieval and vector search cache
5.1 What is it?
RAG systems repeatedly receive similar queries. "What is our leave policy?" might be asked hundreds of times. If your retrieval pipeline runs the full path every time, you're repeating work:
- Query
- Embedding
- Vector DB
- Top-k documents
A retrieval cache stores the result.
5.2 Architecture
- Query
- Normalize
- Cache key
- Retrieval cache
- Hitreturn documentsMissquery the vector database
- Top-k documents
- Store in cache
5.3 Example
retrieval_cache = {}
def retrieve(query, vector_db, top_k=10):
key = f"{query}:{top_k}"
if key in retrieval_cache:
return retrieval_cache[key]
results = vector_db.search(query=query, top_k=top_k)
retrieval_cache[key] = results
return results5.4 But there is a major problem
Documents change.
On Monday, the refund policy says 30 days. You cache "What is the refund policy?" → 30 days. On Tuesday, the policy changes to 14 days. Your cache might still return 30 days.
This is a cache invalidation problem.
5.5 Production solution
Make these part of the cache key:
- the document version
- the index version
- the query
- the retrieval configuration
For example:
key = (
f"{query}:"
f"{index_version}:"
f"{top_k}:"
f"{retrieval_version}"
)Now when the index changes, old entries simply stop matching:
- index_v1invalidated
- index_v2new retrieval
6. Reranking cache
6.1 What is it?
Modern RAG systems often use a two-stage pipeline:
- Retriever
- Top 50 documents
- Reranker
- Top 5 documents
Reranking can be expensive. For a repeated query with the same candidate documents, the reranking result can be cached.
6.2 Architecture
- Query + candidate IDs
- Reranker cache
- Hitreturn ranked resultsMissrun the reranker
- Store in cache
6.3 Cache key
Don't key on the query alone, because the candidate documents might change. Better:
import hashlib
def rerank_key(query, document_ids, model_version):
value = query + ":" + ",".join(document_ids) + ":" + model_version
return hashlib.sha256(value.encode()).hexdigest()This protects against returning a ranking generated from a different document set.
7. Tool result cache
7.1 What is it?
This becomes extremely important in agentic AI.
Imagine an agent with a weather tool, a database tool, a CRM tool, a search tool, a pricing tool and an inventory tool. The agent might repeatedly call get_customer(123) or get_inventory("SKU-123").
If the underlying information doesn't change frequently, caching the result can save:
- latency
- API costs
- database load
- rate limits
7.2 Example
tool_cache = {}
def cached_tool_call(tool_name, arguments, tool_fn):
key = (tool_name, tuple(sorted(arguments.items())))
if key in tool_cache:
return tool_cache[key]
result = tool_fn(**arguments)
tool_cache[key] = result
return resultUsage:
customer = cached_tool_call(
"get_customer",
{"customer_id": "123"},
get_customer,
)7.3 The most important concept: TTL
Not every tool should have the same time to live (TTL). For example:
| Tool | TTL |
|---|---|
| Inventory | 10 seconds |
| Exchange rate | 1 minute |
| Weather | 5 minutes |
| Customer profile | 10 minutes |
| Product details | 1 hour |
| Company policy | 24 hours |
Caching isn't simply cache = true. A cache policy is TTL + freshness + consistency + invalidation.
8. API response cache
8.1 What is it?
AI applications often depend on external APIs: weather, payments, CRM, search and internal services, all called by the LLM or the code around it. Repeated API calls can become expensive or slow. A response cache stores the API response.
8.2 Example
import time
api_cache = {}
def cached_api_call(key, api_fn, ttl=300):
now = time.time()
if key in api_cache:
value, timestamp = api_cache[key]
if now - timestamp < ttl:
return value
result = api_fn()
api_cache[key] = (result, now)
return resultUsage:
weather = cached_api_call(
"weather:bengaluru",
lambda: fetch_weather("bengaluru"),
ttl=300,
)8.3 Production consideration
Never blindly cache every API. Ask:
- Is the response deterministic?
- Can it become stale?
- Does it contain private information?
- Does the API prohibit caching?
- What is the correct TTL?
- Does the user have permission to see it?
9. Conversation and context cache
9.1 What is it?
Conversational agents repeatedly need access to conversation state. Imagine:
User: My order number is 12345.
Assistant: Got it.
User: Where is it now?
The second request needs context from the first. Instead of reconstructing the entire conversation from the database every time, applications often keep a session or context cache.
9.2 Architecture
The conversation ID is the key. The context cache holds:
- recent messages
- user preferences
- the current task
- tool results
- session metadata
9.3 Redis example
import json
import redis
redis_client = redis.Redis(host="localhost", port=6379, decode_responses=True)
def save_conversation(conversation_id, messages):
redis_client.setex(
f"conversation:{conversation_id}",
3600,
json.dumps(messages),
)
def get_conversation(conversation_id):
data = redis_client.get(f"conversation:{conversation_id}")
if not data:
return []
return json.loads(data)9.4 Don't cache the entire conversation forever
Long conversations create another problem:
- 10 messages
- 50 messages
- 500 messages
- 5,000 messages
Eventually the context becomes expensive. Instead of sending the entire history every time, production systems often combine:
- recent messages
- a conversation summary
- important facts
- long-term memory
10. Agent state cache
10.1 What is it?
This is particularly important for agentic RAG. An agent isn't simply input → output. It may run:
- Goal
- Plan
- Step 1tool → result
- Step 2tool → result
- Decision
- Final answer
This intermediate state can be cached.
10.2 Example
Suppose an agent is researching a company. Its task has five steps:
- Search for the company
- Find the financial report
- Extract revenue
- Find competitors
- Generate a summary
If the agent crashes after step 4, you don't want to start again from step 1. Store its state:
{
"task_id": "task_123",
"status": "running",
"completed_steps": [
"company_search",
"financial_report",
"revenue_extraction",
"competitor_search"
],
"results": {
"revenue": "...",
"competitors": ["...", "..."]
}
}Now the agent can resume.
10.3 Simple implementation
agent_state = {}
def save_state(task_id, state):
agent_state[task_id] = state
def load_state(task_id):
return agent_state.get(task_id)In the agent:
state = load_state(task_id)
if not state:
state = {"completed_steps": [], "results": {}}
if "search_company" not in state["completed_steps"]:
result = search_company()
state["results"]["company"] = result
state["completed_steps"].append("search_company")
save_state(task_id, state)This turns an agent from stateless execution into recoverable execution.
11. Putting the 10 caches together
A production RAG + agent architecture might look like this:
- User
- API / gateway
- Semantic cache
- Conversation context cache
- Agent
- RAGToolsAgent state cache
- Embedding cacheTool result cache
- Retrieval cacheRAG path continues
- Reranking cache
- LLM
- Prompt / KV cache
This is why "AI caching" is much bigger than caching the final LLM response.
12. Designing the cache layer
12.1 Cache keys are a first-class design problem
One of the biggest mistakes in production AI systems is a weak cache key.
Bad:
key = queryBetter: include every version the result depends on.
import hashlib
def make_cache_key(
query,
model_version,
prompt_version,
knowledge_version,
retrieval_version,
):
raw = "|".join([
query,
model_version,
prompt_version,
knowledge_version,
retrieval_version,
])
return hashlib.sha256(raw.encode()).hexdigest()This prevents old results from leaking into new system versions.
12.2 Cache invalidation is harder in AI
Traditional software already has the famous problem:
"There are only two hard things in Computer Science: cache invalidation and naming things."
AI systems make invalidation even more complicated. One change ripples through the whole stack:
- Document changes
- Embedding changes
- Vector index changes
- Retrieval changes
- Reranking changes
- Prompt changes
- Model changes
Your cached answer may depend on all of them. So production cache keys should often encode versions:
queryembedding_versionindex_versionreranker_versionprompt_versionmodel_version
12.3 TTL strategy
Don't use one TTL for everything. A better starting point:
| Cache | Typical strategy |
|---|---|
| KV cache | Inference lifecycle |
| Prompt cache | Provider / inference policy |
| Semantic cache | Minutes → hours |
| Embedding cache | Long-lived |
| Retrieval cache | Minutes → hours |
| Reranking cache | Minutes → hours |
| Tool result cache | Seconds → hours |
| API response cache | API-dependent |
| Conversation cache | Session-based |
| Agent state cache | Until task completion |
These are starting points, not universal values. Your actual TTL should be determined by data volatility, correctness requirements, cost, traffic, latency and compliance.
12.4 The security problem
Caching AI results can introduce serious security vulnerabilities. Consider:
- User A"What is my salary?"
- Response cached
- User Basks a similar question
- Cache hit
- User A's salary is returned to User B
That's a catastrophic failure. So cache keys may need:
tenant_iduser_id- permissions
- the query
- the version
For example:
key = (
f"tenant:{tenant_id}:"
f"user:{user_id}:"
f"query:{query_hash}"
)For multi-tenant AI applications:
Never assume a cache is safe to share across users or tenants.
12.5 Cache observability
If you introduce caching without observability, you won't know whether it actually helps. Track:
- cache hit rate and miss rate
- hit latency and miss latency
- saved tokens, saved API calls and saved cost
- stale response rate
- invalidation rate and eviction rate
- cache size
- error rate
A useful dashboard might show:
| Metric | Value |
|---|---|
| Semantic cache hit rate | 68% |
| Embedding cache hit rate | 94% |
| Retrieval cache hit rate | 41% |
| Tool cache hit rate | 73% |
| Tokens saved | 18.4M |
| API calls avoided | 42,891 |
| Estimated cost saved | $1,284 |
| Stale responses | 0.03% |
| Cache errors | 0.01% |
The key metric isn't:
"How many cache hits do we have?"
It's:
"How much useful work did the cache eliminate without hurting correctness?"
12.6 Cache hit rate isn't everything
A 95% cache hit rate sounds great. But if the 5% of misses are the expensive requests, with 10-second latency, huge LLM calls and multiple tool calls, your system can still feel slow.
Track the expected values instead:
Expected latency = hit rate × hit latency + miss rate × miss latency
Expected cost = hit rate × cache cost + miss rate × full pipeline cost12.7 A production cache policy
Instead of:
if cache:
return cachethink:
def get_cached_result(key, ttl, tenant, permissions, version):
# 1. Validate tenant isolation
# 2. Validate permissions
# 3. Validate version
# 4. Check TTL
# 5. Check freshness
# 6. Return cached result
# 7. Otherwise execute pipeline
passCaching is not just a performance feature. It becomes part of your correctness architecture.
13. Deciding what to cache
13.1 A practical decision framework
Before adding a cache, ask five questions.
1. Is the computation expensive?
LLM inference, embeddings, reranking, external APIs and database queries usually are. If the work is cheap, caching may not matter.
2. Is the result reusable?
If every request is unique, the cache hit rate will be close to 0%. Don't cache it.
3. Can the result become stale?
If yes, you must design TTL + versioning + invalidation.
4. Can one user see another user's result?
If yes, you need tenant isolation + authorization-aware keys.
5. Can an incorrect cached result cause harm?
In healthcare, finance, legal, security and enterprise permissions, the answer may be yes. In these cases, cache aggressively only where correctness remains provable.
13.2 What should you cache first?
Don't implement all 10 at once. Start where your system repeats the most work.
- Embedding cache
- Retrieval cache
- Reranking cache
- Semantic cache
- Tool result cache
- Agent state cache
- Semantic cache
- Retrieval cache
- Conversation / context cache
- Prompt cache
- Semantic cache
- Tool result cache
- KV cache
- Prompt / prefix cache
- Distributed cache
13.3 The golden rule
Don't ask:
"Where can I add caching?"
Ask:
"Where am I repeatedly doing expensive work that produces a reusable result?"
That shift in thinking is much more useful. Your AI system might repeatedly perform the same:
- token processing
- prompt processing
- embedding
- retrieval
- reranking
- tool call
- API call
- conversation reconstruction
- agent step
- answer
Each one is a potential caching opportunity.
14. The full picture
14.1 What caching buys you, and what it costs
A mature AI application can eventually run every layer from the architecture above. The result is not simply a faster application. A well-designed cache architecture gives you:
- AI system
- Lower latencyLower costLower load
- Better scale
But caching introduces its own engineering problems:
- invalidation
- freshness
- consistency
- security
- multi-tenancy
- versioning
- memory
- eviction
- observability
That's why production AI caching is not about adding Redis somewhere in the architecture. It is about designing which work can safely be reused, for whom, for how long, and under which version of the system.
14.2 Part 1 takeaway
The 10 core caches are:
- KV cache
- Prompt cache
- Semantic cache
- Embedding cache
- Retrieval / vector search cache
- Reranking cache
- Tool result cache
- API response cache
- Conversation / context cache
- Agent state cache
And they operate at different layers:
| Layer | Caches |
|---|---|
| Model | KV cache, prompt cache |
| RAG | Embedding cache, retrieval cache, reranking cache |
| Agents | Tool result cache, agent state cache |
| Conversation | Context cache, semantic cache |
| Infrastructure | API response cache |
The most important production mindset is simple:
Cache the work, not just the answer.
Because in an AI system, the expensive work often happens before the final answer is generated.
Next: Part 2, Agentic AI caching covers tools, workflows, state, checkpoints, idempotency and MCP.