The AI Caching Playbook, Part 4: RAG Caching
Make retrieval faster, cheaper and smarter: caching queries, embeddings, retrieval, reranking, chunks, context and answers without serving stale data.

On this page
This is Part 4 of the AI Caching Playbook. Part 1 covered the core caches, Part 2 agents and Part 3 conversations. This part goes deep on retrieval.
A RAG system is often described as:
- User query
- Retrieve documents
- Rerank
- Build context
- LLM
- Answer
That looks simple. In production, however, every stage can involve expensive computation:
- Query
- Query preprocessing
- Embedding model
- Vector database
- Keyword search
- Hybrid search
- Reranker
- Document retrieval
- Context construction
- LLM
And many of these operations are repeated unnecessarily:
- The same user may ask the same question again.
- Different users may ask semantically identical questions.
- The same document chunks may be retrieved thousands of times.
- The same query may produce the same rewritten queries.
- The same candidate documents may be reranked repeatedly.
- The same context may be sent to the LLM over and over.
This is where RAG caching becomes a critical production architecture decision.
The goal isn't simply:
"Make RAG faster."
The real goal is:
Retrieve less. Compute less. Send less context. Call the LLM less, without serving stale, unauthorized or incorrect information.
1. Where caching fits inside RAG
A production RAG system can have caching at almost every stage:
- User query
- Query cache
- Query transform cache
- Embedding cache
- Vector search cacheKeyword search cache
- Rerank cache
- Chunk cache
- Context cache
- Response cache
- LLM
Each cache solves a different problem. And each one needs a different cache key.
2. Query cache
2.1 What is it?
The simplest RAG cache is the query cache. Suppose users repeatedly ask:
"What is the leave policy?"
There is no reason to run the entire retrieval pipeline every time if the underlying knowledge hasn't changed. Instead:
- Query
- Normalize
- Hash
- Cache lookup
2.2 Basic implementation
import hashlib
import redis
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
def normalize_query(query: str) -> str:
return " ".join(query.lower().strip().split())
def query_cache_key(query: str) -> str:
normalized = normalize_query(query)
digest = hashlib.sha256(normalized.encode()).hexdigest()
return f"rag:query:{digest}"
def get_cached_query(query: str):
key = query_cache_key(query)
return r.get(key)
def cache_query(query: str, result: str):
key = query_cache_key(query)
r.setex(key, 300, result)2.3 Why the query alone isn't enough
This is not enough for production. "What is the leave policy?" may have different answers for Company A and Company B, for India and the US, or for different employee types.
So hash(query) alone is dangerous. A production key might include tenant + user scope + normalized query + knowledge version + retrieval configuration.
3. Embedding cache
3.1 What is it?
Embedding generation can become surprisingly expensive at scale. At 10,000 queries a day, that's 10,000 embedding API calls, and many of those queries repeat.
Instead of sending every query straight to the embedding model, check a cache first:
- Query
- Embedding cache
- Hitreturn vectorMisscall embedding model
- Store vector in cache
3.2 Important production rule
The embedding cache key should not simply be hash(text). It should include the embedding configuration:
- the text
- the embedding model
- the model version
- the preprocessing version
Otherwise:
- Old embedding model
- Old cached vector
- New embedding model
- Incompatible vector space
- Bad retrieval
3.3 Implementation
import hashlib
import json
import redis
r = redis.Redis(host="localhost", port=6379)
EMBEDDING_MODEL = "embedding-model"
EMBEDDING_VERSION = "v3"
PREPROCESSING_VERSION = "v2"
def embedding_cache_key(text: str) -> str:
normalized = " ".join(text.lower().strip().split())
payload = {
"text": normalized,
"model": EMBEDDING_MODEL,
"model_version": EMBEDDING_VERSION,
"preprocessing": PREPROCESSING_VERSION,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"rag:embedding:{digest}"Then:
def get_embedding(text: str, embedding_model):
key = embedding_cache_key(text)
cached = r.get(key)
if cached:
return json.loads(cached)
vector = embedding_model.embed(text)
r.set(key, json.dumps(vector))
return vectorThis can eliminate a huge number of repeated embedding computations.
4. Vector search cache
4.1 What is it?
Even after caching embeddings, vector search itself still costs computation:
- Query
- Embedding
- Vector DB
- Top 20 chunks
If the same query keeps arriving, the vector database doesn't need to run the search every time. Cache on query embedding fingerprint + index version + top_k + filters.
4.2 Example
import hashlib
import json
def retrieval_cache_key(tenant_id, query, index_version, top_k, filters):
payload = {
"tenant": tenant_id,
"query": normalize_query(query),
"index_version": index_version,
"top_k": top_k,
"filters": filters,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"rag:retrieval:{digest}"Cached result:
{
"chunks": [
{ "id": "chunk_102", "score": 0.91 },
{ "id": "chunk_481", "score": 0.88 }
]
}Notice something important: we are caching chunk IDs and scores, not necessarily the entire document. That gives us more flexibility when document content changes.
5. Retrieval result cache
5.1 What is it?
A useful abstraction is to cache the entire retrieval result:
- Query
- Retriever
- Top K
- Cache
5.2 Example
def get_retrieval_results(query, tenant_id, retriever, index_version, top_k=10):
key = retrieval_cache_key(
tenant_id=tenant_id,
query=query,
index_version=index_version,
top_k=top_k,
filters={},
)
cached = r.get(key)
if cached:
return json.loads(cached)
results = retriever.search(query, top_k=top_k)
r.setex(key, 300, json.dumps(results))
return resultsThis can dramatically reduce vector database load for repeated queries.
5.3 Why the index version belongs in the cache key
Imagine your knowledge base contains Employee Handbook v1. The query "How many vacation days do I get?" returns 20 days, and you cache it.
Then the company updates the handbook to v2, and the correct answer becomes 25 days. If the cache key is only hash(query), the old result can survive.
Instead, the index version moves from index:v1 to index:v2. Now query + index:v1 and query + index:v2 are different cache entries:
- query + index:v1old answer, left to expire
- query + index:v2new answer
This is one of the simplest and safest cache invalidation strategies for RAG.
6. Reranking cache
6.1 What is it?
Many production RAG pipelines look like:
- Vector search
- 50 candidates
- Reranker
- Top 5
Reranking can be expensive. If the same query arrives with the same candidate chunks, the ranking can often be reused.
The cache key should include query + candidate IDs + reranker model + reranker version.
6.2 Implementation
def rerank_cache_key(query, candidate_ids, reranker_model, reranker_version):
payload = {
"query": normalize_query(query),
"candidates": candidate_ids,
"model": reranker_model,
"version": reranker_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"rag:rerank:{digest}"Then:
def rerank(query, candidates, reranker):
candidate_ids = [c["id"] for c in candidates]
key = rerank_cache_key(query, candidate_ids, "reranker", "v2")
cached = r.get(key)
if cached:
return json.loads(cached)
ranked = reranker.rank(query, candidates)
r.setex(key, 600, json.dumps(ranked))
return ranked7. Document and chunk cache
7.1 What is it?
Retrieval might return chunk_102, chunk_481 and chunk_992. But the actual content may live in PostgreSQL, S3, MongoDB or another document store, and fetching the same chunks again and again adds latency.
So cache on chunk ID + document version.
7.2 Example
def chunk_cache_key(document_id, chunk_id, document_version):
return f"rag:chunk:{document_id}:{chunk_id}:{document_version}"Fetching:
def get_chunk(document_id, chunk_id, document_version, database):
key = chunk_cache_key(document_id, chunk_id, document_version)
cached = r.get(key)
if cached:
return json.loads(cached)
chunk = database.get_chunk(document_id, chunk_id)
r.setex(key, 3600, json.dumps(chunk))
return chunk8. Query transformation cache
8.1 What is it?
Modern RAG systems often transform user queries before retrieval. For example, the user asks:
"What changed in our refund policy?"
and the system generates:
- refund policy changes
- updated refund policy
- latest refund rules
- refund policy revisions
That transformation itself costs an LLM call. Cache it, keyed on original query + transformation prompt version + model version.
8.2 Example
def transform_query(query, llm, prompt_version="v2", model="model-v3"):
payload = {
"query": normalize_query(query),
"prompt_version": prompt_version,
"model": model,
}
key = "rag:query_transform:" + hashlib.sha256(
json.dumps(payload, sort_keys=True).encode()
).hexdigest()
cached = r.get(key)
if cached:
return json.loads(cached)
transformed = llm.generate(f"Rewrite the following query for retrieval:\n\n{query}")
result = {"queries": transformed}
r.setex(key, 3600, json.dumps(result))
return resultThis is particularly useful for multi-query RAG, query rewriting, HyDE, question decomposition and query expansion.
9. Hybrid search cache
9.1 What is it?
Many serious RAG systems combine dense (vector) search with BM25 or keyword search:
- Query
- Vector searchKeyword search
- Merge results
You can cache both independently, under keys like dense:{query} and sparse:{query}, and fuse the two cached result sets.
9.2 Why caching them separately helps
This gives an interesting advantage. If you change the fusion algorithm, you don't necessarily need to rerun both searches. You can reuse the cached candidate sets and only redo the fusion:
- Cached dense resultsCached sparse results
- New fusion
10. Context cache
10.1 What is it?
Retrieval is not the end. You still need to assemble:
- system instructions
- conversation context
- the user query
- retrieved documents
- metadata
That assembled context can also be cached. For example:
def context_cache_key(query, retrieved_chunk_ids, prompt_version, index_version):
payload = {
"query": normalize_query(query),
"chunks": retrieved_chunk_ids,
"prompt_version": prompt_version,
"index_version": index_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"rag:context:{digest}"Then:
def build_context(query, chunks, prompt_version, index_version):
chunk_ids = [c["id"] for c in chunks]
key = context_cache_key(query, chunk_ids, prompt_version, index_version)
cached = r.get(key)
if cached:
return cached
context = "\n\n".join(c["text"] for c in chunks)
r.setex(key, 600, context)
return context11. RAG response cache
11.1 What is it?
The final layer is the response cache:
- Query
- RAG
- LLM
- Answer
- Cache
This can provide the biggest latency and cost improvement. But it is also the most dangerous cache to get wrong.
The naive approach is hash(query). That is not enough: the answer depends on much more than the query.
11.2 The RAG response cache key
A production response cache key may contain:
- tenant
- user scope
- query
- conversation context fingerprint
- model version
- prompt version
- knowledge / index version
- retrieval configuration version
11.3 Implementation
def response_cache_key(
tenant_id,
user_scope,
query,
context_fingerprint,
model_version,
prompt_version,
knowledge_version,
retrieval_version,
):
payload = {
"tenant": tenant_id,
"user_scope": user_scope,
"query": normalize_query(query),
"context": context_fingerprint,
"model": model_version,
"prompt": prompt_version,
"knowledge": knowledge_version,
"retrieval": retrieval_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"rag:response:{digest}"Now the response cache becomes much safer.
11.4 Why query-only response caching fails
Imagine two users ask "What is our refund policy?". User A is an enterprise customer and User B is on the free plan. The answer may be different.
Or User A asks "What is my balance?". The answer depends on the user ID, the account ID, permissions and current state.
Therefore:
Same query ≠ same answer.
This is one of the most important principles in RAG caching.
12. Semantic response cache
12.1 What is it?
Exact-match caching catches "What is the refund policy?" but not "Can you explain our refund rules?". These questions may be semantically identical, and a semantic cache can detect that:
- Query
- Embedding
- Vector similarity
- Find the most similar cached query
- Above thresholdreturn cached answerBelow thresholdrun the pipeline
12.2 Example
from numpy import dot
from numpy.linalg import norm
def cosine_similarity(a, b):
return dot(a, b) / (norm(a) * norm(b))
def semantic_cache_lookup(query_embedding, cached_entries, threshold=0.92):
best_match = None
best_score = 0
for entry in cached_entries:
score = cosine_similarity(query_embedding, entry["embedding"])
if score > best_score:
best_score = score
best_match = entry
if best_score >= threshold:
return best_match
return None12.3 Similar is not the same
A semantic cache introduces another correctness problem: semantically similar does not always mean the same answer.
"What was our revenue in 2025?" and "What will our revenue be in 2026?" may be semantically close. They are not interchangeable.
So semantic caching should combine similarity + metadata + freshness + knowledge version + tenant scope.
13. Invalidation
13.1 Index versioning
Index versioning is one of the most powerful techniques in RAG caching. Instead of trying to delete every dependent cache entry by hand when a document changes (potentially 10 million keys), version the knowledge base.
For example, the knowledge base is at knowledge:v101. After ingestion it becomes knowledge:v102. New cache keys automatically use v102, and old entries expire naturally:
KNOWLEDGE_VERSION = "v102"
key = f"rag:response:tenant123:knowledge:{KNOWLEDGE_VERSION}:..."This dramatically simplifies invalidation.
13.2 Document-level invalidation
Sometimes global versioning is too aggressive. If only HR-policy.pdf changes, you don't want to invalidate the entire knowledge base.
You can keep a version per document, such as document:hr-policy:v7 and document:security-policy:v3, and make cache keys depend on the versions of the documents that were actually retrieved:
- Query
- Retrieved chunks
- Document versions
- Cache key
When the HR policy moves from v7 to v8, only the caches that depend on that document become stale.
13.3 Invalidation through dependencies
A production RAG system has a chain of dependencies:
- Document
- Chunks
- Embeddings
- Retrieval
- Reranking
- Context
- Response
When a document changes, every layer after it can become invalid: its chunks, their embeddings, retrieval results, rankings, the assembled context and the final response.
This dependency graph is extremely important. You should always be able to answer:
Which cache entries become invalid when this data changes?
14. Tenants and permissions
14.1 Multi-tenant RAG caching
This is where caching becomes a security problem. Imagine:
- Tenant A"What is the pricing?"
- Answer cached
- Tenant B"What is the pricing?"
- Cache hit
- Tenant A's answer returned to Tenant B
If the cache key is hash(query), Tenant B can receive Tenant A's answer. That's a serious data isolation failure.
Every tenant-sensitive cache should include the tenant scope:
def tenant_cache_key(tenant_id, query, knowledge_version):
payload = {
"tenant": tenant_id,
"query": normalize_query(query),
"knowledge": knowledge_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"tenant:{tenant_id}:rag:{digest}"And tenant isolation must not rely only on the cache key. Authorization should still be enforced before returning protected data.
14.2 ACL and permission-aware caching
Consider a company knowledge base with public, employee, manager, finance and executive documents. Two users can ask "What is the company financial forecast?" but have different permissions.
So the authorization scope may need to be part of the cache key alongside the query and tenant, for example:
tenant_iduser_roledepartmentpermission_scope
But don't cache raw authorization decisions for long periods. Permission changes must invalidate or bypass the affected cache entries.
15. Stampedes and staleness
15.1 Cache stampede
A cache can create another problem. Suppose 100 requests arrive at once and the cache entry doesn't exist yet. All 100 see a miss, and all 100 run embedding, vector search, reranking and the LLM.
- 100 simultaneous misses
- 100 full pipeline runs
You just multiplied the expensive operation. This is called a cache stampede.
15.2 Preventing stampedes with locks
A simple approach is a distributed lock:
import time
def get_or_generate(key, generator, ttl=600):
cached = r.get(key)
if cached:
return cached
lock_key = f"{key}:lock"
acquired = r.set(lock_key, "1", nx=True, ex=10)
if acquired:
try:
cached = r.get(key)
if cached:
return cached
result = generator()
r.setex(key, ttl, result)
return result
finally:
r.delete(lock_key)
for _ in range(20):
time.sleep(0.1)
cached = r.get(key)
if cached:
return cached
return generator()The key principle:
Only one request regenerates. The others wait for its result.
15.3 Stale-while-revalidate
Sometimes you don't need perfect freshness. A company FAQ, product documentation or general knowledge can tolerate a few seconds or minutes of staleness.
Instead of making everyone wait while an expired entry is regenerated, return the stale response and refresh it in the background:
- Request
- Cache
- FreshreturnStalereturn immediately, refresh in background
This is especially useful for high-traffic RAG systems.
16. Deciding what to cache
16.1 What should not be cached?
Caching everything is not a production strategy. Be careful with:
- real-time balances
- personal financial information
- highly dynamic inventory
- current transaction status
- authorization decisions
- sensitive personal data
- time-critical operational information
For dynamic information, the TTL may need to be seconds rather than hours. And for some data, no cache is the correct answer.
16.2 Cache TTL strategy
Different layers should have different TTLs:
| Cache | Example TTL |
|---|---|
| Query transformation | 1 hour |
| Embeddings | Days / weeks |
| Retrieval results | Minutes |
| Reranking | Minutes |
| Chunk content | Hours |
| Context | Minutes |
| FAQ response | Minutes / hours |
| Real-time response | No cache |
| User permissions | Very short / event-driven |
These are starting points, not universal values. TTL should depend on data volatility, correctness requirements, traffic, cost and your freshness SLA.
16.3 Cache cost vs cache value
Caching isn't automatically beneficial.
| Cache lookup | Work it replaces | Benefit |
|---|---|---|
| 5 ms | Vector search, 20 ms | Small |
| 5 ms | LLM generation, 2 seconds | Huge |
So prioritize caches by cost avoided × probability of reuse × correctness confidence. A practical priority order is often:
- Embedding cache
- Retrieval cache
- Reranking cache
- Document / chunk cache
- Response cache
But your actual traffic pattern should determine the order.
17. Putting it together
17.1 A complete production RAG cache
We can now combine everything:
- User query
- Exact query cacheon a miss, continue
- Query transformation
- Embedding cache
- Retrieval cachedense + sparse
- Rerank cache
- Chunk cache
- Context cache
- Response cacheon a miss, call the LLM
- LLM
- Final answer
17.2 A reusable cache key builder
One of the best things you can do in production is centralize cache-key construction. Instead of f"cache:{query}" scattered everywhere, create one abstraction:
import hashlib
import json
class RAGCacheKeyBuilder:
def __init__(self, tenant_id, knowledge_version, model_version, prompt_version):
self.tenant_id = tenant_id
self.knowledge_version = knowledge_version
self.model_version = model_version
self.prompt_version = prompt_version
def _hash(self, payload):
return hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
def embedding(self, text):
return self._hash({
"tenant": self.tenant_id,
"text": normalize_query(text),
"model": self.model_version,
})
def retrieval(self, query, top_k, filters):
return self._hash({
"tenant": self.tenant_id,
"query": normalize_query(query),
"knowledge": self.knowledge_version,
"top_k": top_k,
"filters": filters,
})
def rerank(self, query, candidate_ids):
return self._hash({
"tenant": self.tenant_id,
"query": normalize_query(query),
"candidates": candidate_ids,
"model": self.model_version,
})
def response(self, query, context_fingerprint):
return self._hash({
"tenant": self.tenant_id,
"query": normalize_query(query),
"context": context_fingerprint,
"knowledge": self.knowledge_version,
"model": self.model_version,
"prompt": self.prompt_version,
})Now your application has one place where the cache correctness rules live. That is much safer than scattering cache-key logic across the codebase.
17.3 Production RAG flow with caching
A realistic implementation might look like this:
def rag(query, tenant_id, user, retriever, reranker, llm):
key_builder = RAGCacheKeyBuilder(
tenant_id=tenant_id,
knowledge_version="v102",
model_version="model-v3",
prompt_version="prompt-v5",
)
# 1. Check the final response cache
context_fingerprint = get_context_fingerprint(tenant_id, user)
response_key = key_builder.response(query, context_fingerprint)
cached_response = r.get(response_key)
if cached_response:
return cached_response
# 2. Retrieve documents
retrieval_key = key_builder.retrieval(
query=query,
top_k=20,
filters={"department": user.department},
)
cached_results = r.get(retrieval_key)
if cached_results:
candidates = json.loads(cached_results)
else:
candidates = retriever.search(query, top_k=20)
r.setex(retrieval_key, 300, json.dumps(candidates))
# 3. Rerank
ranked = reranker.rank(query, candidates)
# 4. Build context
context = "\n\n".join(item["text"] for item in ranked[:5])
# 5. Generate the answer
answer = llm.generate(query=query, context=context)
# 6. Cache the response
r.setex(response_key, 300, answer)
return answerThis is intentionally simplified. A real production system would also handle authorization, timeouts, fallbacks, observability, cache stampedes, freshness, index versions, model versions, retry policies, PII and encryption.
18. Observability
18.1 Cache observability
A cache without metrics is just a hidden source of bugs. Track at least hits, misses, hit rate, latency, cache size, evictions, stale reads, lock contention and regeneration count.
For RAG specifically, also track embedding calls avoided, retrieval calls avoided, reranker calls avoided, LLM calls avoided, tokens saved and cost saved. For example:
metrics.increment("rag.embedding.cache_hit")
metrics.increment("rag.retrieval.cache_miss")
metrics.increment("rag.llm.calls_avoided")18.2 Measure hit rate per layer
Don't calculate only an overall cache hit rate. Track each layer independently:
| Layer | Hit rate |
|---|---|
| Embedding cache | 92% |
| Retrieval cache | 71% |
| Reranking cache | 64% |
| Context cache | 51% |
| Response cache | 28% |
This tells you where the system is actually benefiting.
A low response-cache hit rate isn't necessarily bad. A system with a 20% response-cache hit rate, but 95% on embeddings and 80% on retrieval, could still be an excellent architecture: it avoids expensive work even when the final answer can't be reused.
19. RAG cache failure modes
19.1 A new class of bugs
Caching introduces its own failures:
| # | Failure | What happens |
|---|---|---|
| 1 | Wrong tenant | Tenant B receives Tenant A's result |
| 2 | Stale knowledge | A document is updated, but the old answer is returned |
| 3 | Old embedding model | A new model is compared against an old cached vector |
| 4 | Old prompt | Prompt v2 is live, but the response was generated with prompt v1 |
| 5 | Permission mismatch | An employee receives a cached result meant for a manager |
| 6 | Semantic false positive | A similar question returns an incorrect cached answer |
| 7 | Cache stampede | 100 misses become 100 LLM calls |
| 8 | Unbounded cache | Millions of keys cause memory pressure and an eviction storm |
Caching therefore belongs to the correctness architecture, not just the performance architecture.
19.2 A better mental model
Don't think:
"Where can I put Redis?"
Think:
"What expensive computation can I safely reuse?"
| Stage | The question to ask |
|---|---|
| Embedding | Can I reuse this vector? |
| Retrieval | Can I reuse these candidates? |
| Reranking | Can I reuse this ranking? |
| Chunk loading | Can I reuse this content? |
| Context building | Can I reuse this context? |
| LLM generation | Can I safely reuse this answer? |
That is the right way to design RAG caching.
20. Before you ship
20.1 Production checklist
Before shipping RAG caching, verify:
Query
- Query normalization implemented
- Exact query cache evaluated
- Semantic cache evaluated carefully
Embeddings
- Model included in the cache key
- Model version included
- Preprocessing version included
- Vector dimension changes handled
Retrieval
- Tenant included
- Filters included
- Top-K included
- Index version included
- Retrieval configuration versioned
Reranking
- Reranker model version included
- Candidate IDs included
- Ranking configuration versioned
Documents
- Document version tracked
- Chunk cache scoped correctly
- Updates invalidate affected data
Response
- User scope considered
- Conversation context considered
- Model version included
- Prompt version included
- Knowledge version included
Security
- Tenant isolation enforced
- Authorization checked
- Sensitive information handled carefully
- Cache access audited
Reliability
- Cache stampede protection
- TTL strategy
- Stale-while-revalidate where appropriate
- Cache failures fail gracefully
Observability
- Hit rate
- Miss rate
- Latency
- Evictions
- LLM calls avoided
- Tokens saved
- Cost saved
21. The full picture
21.1 The production RAG caching principle
A naive RAG system does all of this on every request:
- Every request
- Embed
- Retrieve
- Rerank
- Build context
- LLM
A production RAG system works down a ladder of questions instead, stopping as soon as it can reuse something:
- Every request
- Can I reuse the answer?yes → return it
- Can I reuse the context?yes → go straight to the LLM
- Can I reuse retrieval?yes → skip the search
- Can I reuse the embedding?yes → skip the embedding call
- Compute only what is necessary
That is the real purpose of RAG caching.
Don't cache blindly. Cache computations whose inputs, dependencies, scope and freshness are well-defined.
And remember:
The best cache is not the one with the highest hit rate. It is the one that removes expensive work without compromising correctness.
21.2 Part 4 final architecture
Putting everything together:
- User
- Response cacheon a miss, continue
- Query transform cache
- Embedding cache
- Retrieval cachedense + sparse
- Reranking cache
- Chunk cache
- Context cache
- LLM
- Answer
The important part isn't having all these caches. The important part is knowing which cache can safely reuse which computation.
What's next?
Part 5: Agentic RAG and agent caching
RAG is only one part of modern AI systems. Agentic systems introduce another level of complexity:
- User
- Agent
- Plan
- Tool
- API
- Database
- Search
- Another agent
- LLM
- Tool
- LLM
- Final answer
Now caching becomes much more interesting. Part 5 will cover:
- agent state, planning, tool result and tool metadata caches
- API response, database query and function result caches
- agent trajectory, sub-agent result, workflow step and partial execution caches
- deterministic tool caching, idempotency, and how retries interact with caches
- agent memory, semantic tool and agent context caches, and agent response caching
- cache-aware tool design and multi-agent cache sharing
- cache invalidation in long-running agents, failure recovery and checkpointing
- observability and cost optimization
The central question becomes:
How do you stop an agent from repeatedly thinking, searching, calling tools and recomputing work it has already done?
And that leads directly into one of the most important production concepts in agentic AI:
Caching the work of the agent, not just the response.
Next: Part 5, Cache the agent's work covers agent state, tools, workflow steps, checkpoints and idempotency.