The AI Caching Playbook, Part 10: The Complete Production Architecture
Bringing it all together: how to design caching as a first-class part of an AI system, from keys and versions to reliability, cost and rollout.

On this page
How to design caching as a first-class part of an AI system.
This is Part 10, the final part of the AI Caching Playbook. The earlier parts covered each layer on its own:
- Part 1: The 10 core caches, from application data to embeddings
- Part 2: Agentic AI caching, covering tools, workflows, state and MCP
- Part 3: Conversational AI caching, covering context, sessions and memory
- Part 4: RAG caching, covering queries, embeddings, retrieval and reranking
- Part 5: Cache the agent's work, covering agent state, tool execution and workflow steps
- Part 6: LLM inference caching, covering prompt, prefix and KV caches
- Part 7: Distributed AI caching, covering sharding, hot keys, stampedes and failure handling
- Part 8: Cache invalidation and correctness, covering TTLs, versioned keys, events and testing
- Part 9: Cache observability and cost, covering latency, tokens, cost, ROI and hit quality
But there's one problem. In a real AI application, these caches don't exist independently. A single production request may pass through several of them:
- User
- API gateway
- Authentication
- Response cache
- Conversation cache
- Query cache
- Embedding cache
- Retrieval cache
- Reranking cache
- Agent / RAG
- Tool cache
- LLM
- Response cache
Now caching becomes an architecture problem. The question is no longer "Should I add Redis?". It's:
What work should be reusable, at which layer, for how long, for whom, and under which version?
That's what this final part is about.
1. The AI request lifecycle
1.1 Start with the request
Consider a typical RAG + agent application, and a user who asks:
"What is our refund policy for enterprise customers?"
That request may become:
- User query
- Authentication
- Tenant resolution
- Conversation state
- Query normalization
- Embedding
- Vector search
- Metadata filtering
- Reranking
- Context construction
- Agent reasoning
- Tool calls
- LLM generation
- Response
Every one of these stages potentially does reusable work. That's where caching comes in.
1.2 The complete architecture
A practical architecture looks like this:
- User
- API gateway
- Auth / tenant
- Cache orchestrator
- Response cacheSession cacheSemantic cache
- RAG / agent
- Query cacheEmbedding cacheAgent state
- Vector retrieval
- Retrieval cache
- Reranker cache
- Context builder
- Tools
- Tool cacheExternal APIs
- LLM
- Prefix / KV cacheResponse cache
- Observability
- CostQualityLatency
The important thing: caching surrounds the AI pipeline. It isn't one component.
2. Cache orchestration and keys
2.1 Cache orchestration
You don't want every service independently calling redis.get(...) with its own rules. That quickly becomes unmanageable. Instead, create a cache abstraction:
from dataclasses import dataclass
from typing import Any, Optional
@dataclass
class CacheResult:
hit: bool
value: Optional[Any]
source: str
class CacheManager:
def __init__(self, backend):
self.backend = backend
def get(self, key: str) -> CacheResult:
value = self.backend.get(key)
if value is None:
return CacheResult(hit=False, value=None, source="miss")
return CacheResult(hit=True, value=value, source="cache")
def set(self, key: str, value: Any, ttl: int):
self.backend.set(key, value, ttl)Now application components don't need to know how the cache is implemented.
2.2 Never use one giant cache key
One of the biggest mistakes is:
key = hash(user_query)Tenant A and Tenant B can both ask "What is the refund policy?". The query is identical; the answer may not be. So the key has to carry more:
key = (
tenant_id,
user_id,
cache_type,
query,
model_version,
knowledge_version,
)2.3 A central cache key builder
Instead of building keys by hand everywhere:
import hashlib
import json
class CacheKeyBuilder:
def build(self, cache_type, tenant_id, payload, version):
raw = json.dumps(
{
"cache_type": cache_type,
"tenant_id": tenant_id,
"payload": payload,
"version": version,
},
sort_keys=True,
)
digest = hashlib.sha256(raw.encode()).hexdigest()
return f"ai:{cache_type}:v{version}:{digest}"Usage:
key_builder = CacheKeyBuilder()
key = key_builder.build(
cache_type="response",
tenant_id="tenant_123",
payload={"query": "refund policy"},
version=3,
)Centralizing key generation prevents subtle inconsistencies.
2.4 Cache keys are contracts
A cache key isn't merely an identifier. It defines when two requests are equivalent:
Same query + same tenant + same model + same prompt + same knowledge + same retrieval configuration = potentially the same answer
If any one dependency changes, the cache should usually stop matching.
2.5 Version everything
Production AI systems should version the model, the prompt, the embedding model and its preprocessing, the vector index, documents, the reranker, tools, the agent workflow, schemas, business rules and the cache key format itself. For example:
VERSIONS = {
"prompt": "v7",
"model": "model-x-v3",
"embedding": "embed-v4",
"index": "index-2026-10-08",
"reranker": "rerank-v2",
"workflow": "agent-v5",
}Now cache correctness becomes much easier.
3. The RAG cache pipeline
3.1 Layer by layer
A production RAG request can look like:
- Query
- Normalize
- Query cacheon a miss, continue
- Embedding cacheon a miss, call the embedding model
- Retrieval cacheon a miss, query the vector DB
- Reranking cacheon a miss, run the reranker
- Context cacheon a miss, build the context
- LLM
- Response cache
Each layer has a different responsibility.
3.2 Query cache
Query transformations are often deterministic. "What are the refund terms?" might become "enterprise customer refund policy". Cache it:
query_key = key_builder.build(
cache_type="query",
tenant_id=tenant_id,
payload={
"query": normalized_query,
"transform_version": "v2",
},
version=1,
)3.3 Embedding cache
The embedding cache key must include the embedding model:
embedding_key = key_builder.build(
cache_type="embedding",
tenant_id=tenant_id,
payload={
"text": normalized_query,
"embedding_model": "embedding-v4",
"preprocessing_version": "v3",
},
version=1,
)Don't use hash(text) alone: different models produce different vector spaces.
3.4 Retrieval cache
Retrieval depends on more than the query. A robust retrieval key may include the tenant, a query fingerprint, the embedding model, the index version, top_k, filters, the ACL scope and the retrieval algorithm:
retrieval_key = key_builder.build(
cache_type="retrieval",
tenant_id=tenant_id,
payload={
"query_hash": query_hash,
"index_version": "index-2026-10-08",
"top_k": 10,
"filters": {"department": "finance"},
},
version=2,
)3.5 Reranking cache
Reranking depends on the query + candidate documents + reranker model + reranking configuration:
rerank_key = key_builder.build(
cache_type="rerank",
tenant_id=tenant_id,
payload={
"query": query,
"candidate_ids": candidate_ids,
"reranker": "rerank-v2",
},
version=1,
)3.6 Context cache
Building the context can itself be expensive:
- Retrieve 20 chunks
- Deduplicate
- Rerank
- Compress
- Format
- Apply token budget
Cache the resulting context package:
context_key = key_builder.build(
cache_type="context",
tenant_id=tenant_id,
payload={
"chunk_ids": chunk_ids,
"context_builder_version": "v4",
"token_budget": 8000,
},
version=1,
)3.7 Response cache
Response caching is the most powerful layer, and the most dangerous. Its key should consider the tenant, query, conversation context, model, prompt, knowledge version, retrieval configuration, tool results and agent state:
response_key = key_builder.build(
cache_type="response",
tenant_id=tenant_id,
payload={
"query": query,
"context_fingerprint": context_hash,
"model": "model-x",
"prompt_version": "v7",
"knowledge_version": "v12",
},
version=3,
)3.8 Why the response cache is the riskiest layer
The response cache sits on top of everything else: response cache → retrieval cache → embedding cache. When it hits, you skip everything below it. That's great.
But when it's wrong, you skip everything below it too. That's why response caching needs the strongest correctness guarantees.
4. Conversations, agents and tools
4.1 Conversation caching
Conversational systems add users, conversations, messages, summaries, memory, preferences and the current task. You shouldn't rebuild all of it from scratch on every turn. A session cache might hold:
session_state = {
"conversation_id": "conv_123",
"summary": "...",
"recent_messages": [...],
"memory_version": 8,
}Cache it separately from the final response.
4.2 The conversation cache key
Don't key only on conversation_id, because the conversation keeps changing. Include the message sequence and memory version:
conversation_key = (tenant_id, conversation_id, message_sequence, memory_version)
key = f"conversation:{tenant_id}:{conversation_id}:{message_sequence}"Every new message advances the sequence.
4.3 Agent state cache
Agents are different. They may plan, search, analyze, call tools, validate and generate:
- Plan
- Search
- Analyze
- Tool
- Validate
- Generate
You don't want to lose all of that if the process crashes. Store the state:
agent_state = {
"run_id": "run_123",
"step": 4,
"plan": plan,
"completed_steps": completed_steps,
"pending_steps": pending_steps,
}This enables recovery.
4.4 Checkpointing vs caching
These are related but different:
| Cache | Checkpoint | |
|---|---|---|
| Purpose | Optional optimization | Execution recovery state |
| If deleted | The system can recompute | You may lose progress |
Don't treat critical agent state as disposable cache data.
4.5 Tool caching
Suppose your agent calls get_customer_profile(). The result might be cacheable:
tool_key = key_builder.build(
cache_type="tool",
tenant_id=tenant_id,
payload={
"tool": "get_customer_profile",
"customer_id": customer_id,
"tool_version": "v4",
},
version=1,
)But the cache policy depends on the tool.
4.6 Read tools vs write tools
| Tool type | Examples | Caching |
|---|---|---|
| Read | get_customer(), search_inventory(), get_weather(), fetch_document() |
Usually cacheable |
| Write | create_payment(), send_email(), delete_customer(), create_order() |
Usually not safe as a normal result cache; use an idempotency key |
4.7 Idempotency keys
For example:
idempotency_key = f"payment:{tenant_id}:{order_id}:{operation_id}"The payment service guarantees that the same operation produces the same effect. That is different from "cache hit, so don't execute".
5. Inference-level caching
5.1 Prefix cache
Now move down to the LLM. Consider a prompt made of a system prompt, company policy, tool definitions and the user query. The first three sections may stay the same across thousands of requests, which makes them a perfect candidate for prefix reuse:
| Section | |
|---|---|
| System prompt | Stable prefix |
| Company policy | Stable prefix |
| Tool definitions | Stable prefix |
| User query | Dynamic |
The more stable the prefix, the more reuse you get.
5.2 Cache-friendly prompt design
- System
- Today's timestamp
- Random request ID
- Company policy
- Tool definitions
- User query
- System
- Company policy
- Tool definitions
- Dynamic timestamp
- User query
5.3 KV cache
At the inference layer, the prompt is prefilled into K/V tensors, and each decoded token reuses the previous K/V:
- Prompt
- Prefill
- K/V tensors
- Decode token
- Reuse previous K/V
The KV cache avoids recomputing attention over previous tokens during autoregressive generation:
| Token | Without KV cache | With KV cache |
|---|---|---|
| 1 | Compute | Compute and store |
| 2 | Recompute 1 + 2 | Reuse 1 |
| 3 | Recompute 1 + 2 + 3 | Reuse 1 + 2 |
This is an inference-level optimization. It is different from an application response cache.
5.4 Paged KV memory
Large-scale inference needs careful GPU memory management. A runtime may organize KV cache memory into blocks, or pages:
- KV1KV2KV3KV4…
This allows better sharing and memory utilization across concurrent requests. But remember:
Paged attention is a memory-management technique around the KV cache, not a separate application cache.
6. Lifetimes, ownership and dependencies
6.1 Cache layers have different lifetimes
This is critical:
| Cache | Typical lifetime |
|---|---|
| KV cache | Milliseconds / the request |
| Response cache | Seconds / minutes |
| Session cache | Minutes / hours |
| Embedding cache | Hours / days |
| Document cache | Days / weeks |
| Static knowledge cache | Potentially months |
Don't give every cache a 24-hour TTL. That's lazy architecture.
6.2 Cache ownership
Every cache should have an owner:
| Cache | Owner |
|---|---|
| Response cache | AI platform |
| Embedding cache | RAG platform |
| Tool cache | Tool service |
| Session cache | Conversation service |
| KV cache | Inference runtime |
Without ownership, nobody knows who invalidates it, who changes its schema, who monitors it or who handles incidents.
6.3 The cache dependency graph
Your caches form a dependency graph:
- Documents
- Embedding
- Vector index
- Retrieval
- Reranking
- Context
- LLM
- Response
When documents change, everything downstream is affected. But you don't necessarily need to delete every cache immediately: versioning can invalidate downstream results naturally.
6.4 Version-based invalidation
Instead of redis.delete("millions-of-keys"), bump the version. With index_version = v13, old keys look like retrieval:v12:… and new ones like retrieval:v13:…. Old entries expire naturally.
This dramatically reduces invalidation complexity.
6.5 A version registry
Create one source of truth for versions:
class VersionRegistry:
def __init__(self):
self.versions = {
"prompt": "v7",
"model": "v3",
"embedding": "v4",
"index": "v13",
"reranker": "v2",
"workflow": "v5",
}
def get(self, name):
return self.versions[name]
def update(self, name, version):
self.versions[name] = versionCache key builders then read from this registry.
7. Security
7.1 Multi-tenant caching
In SaaS AI systems, tenants may have completely different knowledge. Never create key = hash(query). Instead:
key = (tenant_id, query, knowledge_version)Tenant isolation must be part of the architecture.
7.2 Authorization is not caching
This distinction is extremely important. Suppose User A can access document_123 and User B cannot. A cached retrieval result containing document_123 must not be served to User B, even if the query is identical.
Authorization must still be enforced. Never let a cache hit become an authorization bypass.
7.3 Cache security boundaries
Think of scopes, from widest to narrowest:
- Public cache
- Tenant cache
- User cache
- Session cache
The more sensitive the data, the narrower the scope should be:
| Data | Scope |
|---|---|
| Public FAQ | Shared cache |
| Tenant documentation | Tenant-scoped cache |
| Personal financial data | User- or session-scoped cache |
8. Reliability
8.1 Cache stampede protection
When a popular entry expires, one miss can become 10,000 simultaneous misses. Instead of letting every request regenerate the value, let one request take a lock and regenerate, while the others wait or reuse the result:
- Lock
- Regenerate
- Populate cache
- Wait, then reuse the result
A simple Redis-style lock:
lock_key = f"lock:{cache_key}"
acquired = redis.set(lock_key, "1", nx=True, ex=10)
if acquired:
# regenerate
value = generate()
redis.set(cache_key, value, ex=300)
redis.delete(lock_key)
else:
# wait / retry / serve stale
passProduction implementations need careful failure handling, but the principle is straightforward.
8.2 Stale-while-revalidate
Another strategy: if the cached entry has expired but is still usable, return it and refresh it in the background.
- Cache entry exists
- Expired but usable
- Return the stale result
- Refresh asynchronously
if cached and cached.age < stale_limit:
trigger_background_refresh()
return cached.valueThis is especially useful when the freshness requirement is flexible and latency is critical.
8.3 Cache failure shouldn't mean AI failure
Suppose Redis goes down.
- Redis down
- Application down
- Redis down
- Cache disabled
- Normal AI pipeline
Caching should generally be a performance layer, not your only source of truth:
try:
cached = cache.get(key)
except Exception:
cached = None
if cached:
return cached
return generate()For critical checkpoint state, though, losing the store may need a different recovery strategy.
8.4 Timeouts matter
Never let a cache lookup hang the whole request:
value = cache.get(key, timeout=0.05)With a 50 ms cache timeout, if the cache doesn't respond, the request continues without it. A 2-second cache timeout defeats the purpose of caching.
8.5 A cache circuit breaker
If your cache backend keeps failing, open a circuit:
- Application
- Circuit breaker
- Closeduse the cacheOpenskip the cache
After recovery, the breaker moves back step by step:
- Open
- Half-open
- Test
- Closed
This protects the AI application from cache infrastructure failures.
9. Warming, admission and eviction
9.1 Cache warming
For predictable workloads, warm the cache: popular FAQ answers, popular documentation, popular product queries, common embeddings and common system prefixes. At deployment:
for query in popular_queries:
response = generate(query)
cache.set(make_key(query), response, ttl=3600)But don't blindly warm millions of entries. Warm only data with predictable demand.
9.2 Precompute expensive work
Sometimes the best cache isn't a runtime cache at all. If you know your 100 most common enterprise questions, don't wait for the first user to pay for each one. Generate the answers during deployment, validate them and cache them.
This turns runtime latency into deployment-time work.
9.3 Cache admission
Not every result should enter the cache. A useful admission policy considers frequency, generation cost, result size, freshness and reuse probability:
def should_cache(reuse_probability, generation_cost, size_mb, freshness):
return (
reuse_probability > 0.3
and generation_cost > 0.01
and size_mb < 5
and freshness > 300
)Again, the thresholds are workload-specific.
9.4 Cache eviction
Common policies include LRU, LFU, TTL, size-based, priority-based and cost-aware eviction. For AI workloads, cost-aware policies are interesting:
| Uses | Regeneration cost | |
|---|---|---|
| Entry A | 100 | $0.001 |
| Entry B | 5 | $1.00 |
Entry B might still deserve its place, because each miss is expensive. So consider:
Expected savings = reuse probability × regeneration cost
9.5 A cost-aware admission score
For example:
def cache_score(reuse_probability, regeneration_cost, memory_cost):
return reuse_probability * regeneration_cost - memory_costCache the entries with high expected value. That is much smarter than caching everything.
10. Observability
10.1 Trace the whole architecture
Every request should produce a trace like this:
| Stage | Result | Time |
|---|---|---|
| Response cache | Miss | 3 ms |
| Session cache | Hit | 2 ms |
| Query cache | Miss | 2 ms |
| Embedding cache | Hit | 1 ms |
| Retrieval cache | Miss | 4 ms |
| Vector DB | 72 ms | |
| Reranker cache | Miss | 2 ms |
| Reranker | 85 ms | |
| Tool cache | Hit | 3 ms |
| LLM | 820 ms | |
| Total | 994 ms |
Now you know exactly where the time went.
10.2 A cache decision trace
A useful debugging record:
trace = {
"request_id": request_id,
"cache": {
"response": "MISS",
"session": "HIT",
"query": "MISS",
"embedding": "HIT",
"retrieval": "MISS",
"reranker": "MISS",
"tool": "HIT",
},
"versions": {
"model": "v3",
"prompt": "v7",
"index": "v13",
},
}When a user asks "Why was this answer generated?", you have the evidence.
11. What to cache
11.1 The golden rule: cache work, not just data
Think about the AI pipeline:
- Raw data
- Transformation
- Embedding
- Retrieval
- Reranking
- Context
- Reasoning
- Tool calls
- Generation
Each stage produces something reusable. So:
Cache the most expensive deterministic work that is likely to be reused.
Not everything.
11.2 What should you cache?
A useful decision table:
| Layer | Usually cache? | Typical scope |
|---|---|---|
| Static documents | Yes | Long |
| Embeddings | Yes | Long |
| Query transformations | Yes | Medium |
| Retrieval results | Yes | Short / medium |
| Reranking | Yes | Short / medium |
| Context | Yes | Short |
| LLM response | Sometimes | Short / medium |
| Session state | Yes | Session |
| Agent checkpoints | Yes, but durably | Workflow |
| Read-only tool results | Often | Tool-specific |
| Write operations | Usually no | Use idempotency |
| KV cache | Runtime-managed | Request / session |
| Prefix cache | Often | Runtime / provider-specific |
11.3 What should you not cache blindly?
Avoid blindly caching real-time financial balances, payment state, authorization decisions, highly dynamic inventory, sensitive personal information, non-idempotent write operations, security decisions and frequently changing data.
The right answer is workload-specific, but the risk should always be explicit.
12. A request, end to end
12.1 The complete request flow
Let's combine everything. A user asks:
"What is our enterprise refund policy?"
| Step | Layer | Result | What happens |
|---|---|---|---|
| 1 | Authentication | User resolved to Tenant A | |
| 2 | Response cache | Miss | |
| 3 | Conversation cache | Hit | |
| 4 | Query cache | Miss | Normalize the query |
| 5 | Embedding cache | Hit | The embedding already exists |
| 6 | Retrieval cache | Miss | Search the vector database |
| 7 | Reranker cache | Miss | Run the reranker |
| 8 | Context cache | Miss | Build the context |
| 9 | Agent | Decide whether tools are needed | |
| 10 | Tool cache | Hit | Reuse existing customer policy metadata |
| 11 | LLM | Generate the answer | |
| 12 | Response cache | Store | Cache the result and return it to the user |
The next identical request may then be just:
- User
- Response cache hit
- Answer
instead of repeating the entire pipeline.
12.2 The latency transformation
| Stage | Without caching | Exact response-cache hit | Lower-level caches only |
|---|---|---|---|
| Authentication | 10 ms | 10 ms | 10 ms |
| Response cache | 5 ms | ||
| Embedding | 30 ms | 2 ms (cached) | |
| Retrieval | 80 ms | 3 ms (cached) | |
| Reranking | 100 ms | 90 ms | |
| Tool calls | 200 ms | ||
| LLM | 1,500 ms | 900 ms | |
| Total | 1,920 ms | 15 ms | 1,005 ms |
An exact response-cache hit takes the request from 1,920 ms to 15 ms. Even when only the lower-level caches hit, you still save almost half the time. This is why layered caching matters.
12.3 The cost transformation
Suppose an uncached request costs:
| Stage | Cost |
|---|---|
| Embedding | $0.0003 |
| Retrieval | $0.0005 |
| Reranker | $0.001 |
| Tools | $0.005 |
| LLM | $0.04 |
| Total | $0.0468 |
A response cache hit might cost about $0.00001 for the lookup, saving $0.04679 per request. At 1,000,000 requests a month, that's roughly $46,790 a month.
But only if the cached answer is correct and fresh enough.
13. Rolling it out
13.1 The real optimization strategy
Don't start with "Where can I add Redis?". Work through these questions in order:
- What is expensive?
- What repeats?
- What is deterministic?
- What can safely be reused?
- What freshness is acceptable?
- How do we invalidate it?
- How do we measure the value?
That is the production caching mindset.
13.2 A practical implementation strategy
Don't implement everything at once:
| Phase | Focus | What to do |
|---|---|---|
| 1 | Measure | Instrument LLM calls, tokens, latency, cost and repeated queries |
| 2 | Cheap caches | Embedding cache, query cache, read-only tool cache |
| 3 | Retrieval | Retrieval results, reranking |
| 4 | Conversations | Session, summary, memory |
| 5 | Responses | Response caching, only after correctness is understood |
| 6 | Inference | Prefix caching, KV cache, batching |
| 7 | Economics | Admission, eviction, TTLs, ROI |
14. Configuration and code
14.1 Production cache configuration
A configuration-driven approach is much easier to manage:
CACHE_POLICIES = {
"embedding": {"ttl": 86400 * 30, "scope": "tenant"},
"retrieval": {"ttl": 3600, "scope": "tenant"},
"reranking": {"ttl": 1800, "scope": "tenant"},
"response": {"ttl": 300, "scope": "tenant"},
"session": {"ttl": 3600, "scope": "user"},
"tool": {"ttl": 300, "scope": "tool-defined"},
}Now cache policy isn't buried throughout the application code.
14.2 A cache policy object
You can make policies explicit and typed:
from dataclasses import dataclass
@dataclass
class CachePolicy:
ttl: int
enabled: bool
scope: str
stale_while_revalidate: bool
max_size_bytes: int
response_policy = CachePolicy(
ttl=300,
enabled=True,
scope="tenant",
stale_while_revalidate=True,
max_size_bytes=500_000,
)14.3 The cache-aside pattern
For many AI workloads:
def get_or_generate(cache, key, generator, ttl):
cached = cache.get(key)
if cached is not None:
return cached
value = generator()
cache.set(key, value, ttl)
return valueSimple and powerful. But the generator must be safe to run concurrently, and for high-traffic keys you should combine this with stampede protection.
14.4 A production cache wrapper
The abstraction can grow to apply policies and record metrics:
class AICache:
def __init__(self, backend, metrics, policies):
self.backend = backend
self.metrics = metrics
self.policies = policies
def get(self, cache_type, key):
policy = self.policies[cache_type]
if not policy.enabled:
return None
value = self.backend.get(key)
self.metrics.record_get(cache_type=cache_type, hit=value is not None)
return value
def set(self, cache_type, key, value):
policy = self.policies[cache_type]
if not policy.enabled:
return
self.backend.set(key, value, ttl=policy.ttl)Now the whole AI application shares one consistent caching interface.
14.5 Observable by default
Never call cache.get(key) without knowing:
- Did it hit?
- How long did it take?
- What did it save?
- Was it fresh?
- Which version was used?
Every cache operation should produce telemetry.
15. The final architecture
15.1 What a mature AI caching layer looks like
- User
- API gateway
- Auth / tenant
- Cache orchestrator
- Response cacheConversation cacheSemantic cache
- RAG / agent
- Query cacheEmbedding cacheAgent state
- Vector search
- Retrieval cache
- Reranking cache
- Context builder
- Tools
- Tool cacheExternal API
- LLM
- Prefix / KV cacheResponse cache
- Observability
- CostQualityLatency
- Cache control
- InvalidationTTL / policyVersions
Observability feeds cache control, and cache control (invalidation, policies and versions) shapes every layer above it. This is what a mature AI caching layer starts to look like.
16. Before you ship
16.1 Production checklist
Before shipping an AI caching architecture:
Architecture
- Cache layers are clearly defined
- Each cache has an owner
- Cache failures don't unnecessarily break the AI system
- Critical state is separate from disposable cache
Keys
- Tenant included
- User and session scope considered
- Model version included
- Prompt version included
- Knowledge and index version included
- Tool version included
- Preprocessing version included
- Retrieval parameters included
Correctness
- TTLs defined
- Freshness requirements defined
- Invalidation strategy defined
- Authorization checked independently
- Semantic cache evaluated
- Versioning implemented
Reliability
- Stampede protection
- Timeouts
- Circuit breaker
- Fallback path
- Hot-key protection
- Eviction policy
Economics
- Tokens saved
- Calls avoided
- Latency saved
- Cost saved
- Cache infrastructure cost
- ROI
Observability
- Hit and miss metrics
- P95 / P99 lookup latency
- Staleness metrics
- Quality metrics
- Cache memory
- Eviction metrics
- Cost dashboard
- Alerts
17. The 10 principles of production AI caching
17.1 The principles
- Cache work, not everything. Cache expensive, reusable computation.
- Keys define correctness. A bad key can create a wrong answer.
- Version your dependencies. Models, prompts, embeddings, indexes and tools change.
- Scope your caches. Tenant, user, session and public data are different.
- Separate the cache from the source of truth. The cache should normally be disposable.
- Don't cache writes blindly. Use idempotency for side effects.
- Protect against stampedes. One expired hot key shouldn't trigger thousands of expensive executions.
- Measure useful work avoided. Tokens, LLM calls, tools, latency and cost matter.
- Evaluate cache quality. A cache hit that returns the wrong answer is not a successful hit.
- Optimize economically. The goal is lower cost and lower latency with the same or better quality.
18. The full picture
18.1 The complete mental model
The entire series reduces to one loop:
- AI request
- Can we reuse this work?
- Yesvalidate version, scope and freshness, then return the resultNocompute the work, then cache it
- Measure value
- CostLatencyQuality
- Optimize
That's the production caching loop.
18.2 The AI Caching Playbook, complete
We've gone from individual cache concepts to a complete production architecture:
- Part 1: The 10 core caches
- Part 2: Agentic AI caching
- Part 3: Conversational AI caching
- Part 4: RAG caching
- Part 5: Cache the agent's work
- Part 6: LLM inference caching
- Part 7: Distributed AI caching
- Part 8: Cache invalidation and correctness
- Part 9: Cache observability and cost
- Part 10: The complete production architecture (this part)
And the whole journey comes down to changing the question. Don't ask "Can I cache this?". Ask:
- What expensive work repeats in my system?
- Can it be safely reused?
- What makes two results equivalent?
- When does it become stale?
- How do I invalidate it?
- How much work does it save?
- Is the quality still correct?
- Is the cache worth its cost?
18.3 The final principle
Caching is not a performance feature you bolt onto an AI application. It is a system-design decision that determines what work your AI system is allowed to reuse.
A production-grade AI system isn't simply an LLM + a vector DB + Redis. It has identity, versioning, reusable work, cache policy, invalidation, freshness, security, observability and economics.
And when all of those work together:
- Same quality
- Lower latencyLower cost
- Less compute
- More throughput
- Better AI system
That's production AI caching. Not just making requests faster: making expensive intelligence reusable.