The AI Caching Playbook, Part 8: Cache Invalidation and Correctness
Fast is useless if the answer is wrong: TTLs, versioned keys, dependency graphs, events, concurrency and testing for correct AI caches.

On this page
Fast is useless if the answer is wrong.
This is Part 8 of the AI Caching Playbook. The earlier parts:
- Part 1: The 10 core caches, from application data to embeddings
- Part 2: Agentic AI caching, covering tools, workflows, state and MCP
- Part 3: Conversational AI caching, covering context, sessions and memory
- Part 4: RAG caching, covering queries, embeddings, retrieval and reranking
- Part 5: Cache the agent's work, covering agent state, tool execution and workflow steps
- Part 6: LLM inference caching, covering prompt, prefix and KV caches
- Part 7: Distributed AI caching, covering sharding, hot keys, stampedes and failure handling
Caching sounds simple:
- Request
- Cache
- Hit
- Return result
But production systems aren't static. Data changes. Documents get updated. Users change permissions. Models are upgraded. Prompts change. Embeddings are regenerated. Vector indexes are rebuilt. Tools return different results. Agents create new state.
And suddenly the cache holds an answer that was correct five minutes ago. That is the fundamental caching problem:
How do you reuse expensive computation without reusing information that is no longer valid?
This matters especially in AI systems. Consider a RAG application:
- Document v1
- Embedding v1
- Vector index v1
- Retrieval
- Reranking
- Context
- LLM
- Answer
Now the document changes to v2. What happens to the embedding, the vector index, the retrieval cache, the reranking cache, the context cache and the response cache?
If you don't answer that question explicitly, your cache architecture is incomplete.
1. The cache correctness problem
1.1 Stale data
Suppose your application caches user:123 → Premium. Then the user downgrades from Premium to Free, but the cache still says Premium. The application may keep treating them as a premium user.
That's stale data. Now imagine the same problem with AI:
- Knowledge base
- Old information
- Cached answer
- User receives an outdated answer
The model isn't necessarily wrong. The cache is.
1.2 Cache invalidation
Cache invalidation means:
Making cached data unusable when its underlying dependencies change.
There are several major approaches: TTLs, explicit deletion, versioning, event-driven invalidation, dependency-based invalidation, namespace invalidation, generational invalidation, write-through and write-behind.
No single strategy works for everything.
2. TTL
2.1 Time to live
The simplest strategy:
redis.set(key, value, ex=300)The entry expires after 300 seconds, which gives you maximum staleness ≈ TTL.
But there's an important catch: a TTL doesn't know whether your data changed. With a 1-hour TTL, if the document changes 10 seconds after caching, the cache can stay stale for almost 59 minutes and 50 seconds.
TTL controls lifetime, not correctness.
2.2 When TTL is enough
A TTL works well when stale data is acceptable: weather, popular content, recommendations, analytics, non-critical metadata, public dashboards. For example:
redis.set("trending_topics", data, ex=60)A one-minute-old trending list may be perfectly fine.
2.3 When TTL is not enough
Be careful with authorization, permissions, financial state, inventory, private data, compliance information, security configuration and critical business rules. A five-minute-stale value might be unacceptable.
For these systems, a TTL should not be your only correctness mechanism.
3. Explicit and dependency-aware invalidation
3.1 Explicit invalidation
The next strategy is direct deletion. When document:123 changes:
redis.delete("document:123")Simple. But AI systems have dependencies. That document may have produced embedding:123, retrieval:abc, rerank:xyz, context:qwe and response:123. Deleting only document:123 doesn't invalidate the derived results.
3.2 Dependency graphs
Think of cached computation as a graph:
- Document
- Embedding
- Vector index
- Retrieval
- Reranking
- Context
- Prompt
- Response
If the document changes, everything downstream may become invalid. This is a dependency problem.
3.3 Dependency-aware invalidation
You can track dependencies explicitly:
dependencies = {
"response:abc": ["context:xyz", "model:v5", "prompt:v12"],
"context:xyz": ["retrieval:qwe"],
"retrieval:qwe": ["document:123", "document:456"],
}
def invalidate_document(document_id):
affected = find_dependents(document_id)
for key in affected:
redis.delete(key)This gives you precise invalidation. The problem: at large scale, maintaining dependency graphs can get expensive.
4. Version-based invalidation
4.1 Change the version, not the keys
This is one of the most useful patterns for AI. Instead of deleting 10 million cache entries, change the version: knowledge:v1 becomes knowledge:v2, and the version is part of every key:
key = (
f"rag:"
f"knowledge:{knowledge_version}:"
f"query:{query_hash}"
)Old entries look like rag:knowledge:1:abc and new ones like rag:knowledge:2:abc. The old cache simply expires naturally.
4.2 Why versioning works so well for AI
AI systems have many versioned dependencies: the model, prompt, embedding model, index, reranker, tools, schema and knowledge base. Instead of trying to understand every invalidation relationship:
- A dependency changes
- Its version changes
- New cache namespace
This is much easier to reason about.
4.3 Embedding versioning
Suppose you start with text-embedding-model-v1 and migrate to text-embedding-model-v2. An embedding produced by v1 must never be mistaken for a v2 embedding. A good key:
key = (
f"embedding:"
f"model:{model_name}:"
f"version:{model_version}:"
f"preprocess:{preprocess_version}:"
f"text:{text_hash}"
)Now changing the model automatically creates a new cache namespace.
4.4 Retrieval versioning
Suppose your vector index moves from index:v10 to index:v11. The retrieval cache should include the index version:
key = (
f"retrieval:"
f"index:{index_version}:"
f"query:{query_hash}:"
f"topk:{top_k}"
)Without index versioning:
- New index
- Old retrieval result
- Incorrect evidence
That's a subtle but serious RAG bug.
4.5 Response cache versioning
A final AI response can depend on many things:
key = (
f"response:"
f"tenant:{tenant_id}:"
f"model:{model_version}:"
f"prompt:{prompt_version}:"
f"knowledge:{knowledge_version}:"
f"index:{index_version}:"
f"query:{query_hash}"
)Now a change to the model, prompt, knowledge or index automatically creates a new response cache.
5. The cache key is a correctness contract
5.1 What a key really says
A cache key isn't just a lookup identifier. It defines:
Under what conditions is this cached result equivalent to another request?
response:abc123 doesn't tell us enough. This does:
response:tenant:acme:model:v5:prompt:v12:knowledge:v8:index:v21:query:abc123Every component answers a correctness question.
5.2 What should go into an AI cache key?
Depending on the cache:
| Area | Possible key components |
|---|---|
| Scope | Tenant, user scope |
| Input | Query, normalized query, filters, top-K |
| Model | Model, model version, temperature and other relevant generation settings |
| Prompt | Prompt version |
| Retrieval | Embedding model and version, preprocessing version, index version, reranker version, knowledge version |
| Tools | Tool version, schema version |
Don't blindly include everything. Only include values that actually affect the cached output.
5.3 Cache-key explosion
There is a downside. If you include temperature, top-K, model, prompt, tenant, user, region and language, the same question can create hundreds of different cache keys, and your hit rate collapses.
So the goal isn't to put everything into the key. The goal is to put every correctness-relevant dependency into the key.
6. Semantic equivalence
6.1 Similar or the same?
"What is the refund policy?" and "Can you tell me your refund policy?" are semantically similar. Should they share a cache entry?
For an exact cache: no. For a semantic cache: potentially. But semantic caching introduces another correctness problem: two questions can be similar without being equivalent.
6.2 Semantic cache risk
"What is the price of Product A?" and "What was the price of Product A last year?" may have very high embedding similarity, but the answers are completely different.
So semantic similarity ≠ semantic equivalence. This is why semantic response caches need conservative thresholds and often extra validation.
6.3 Exact cache vs semantic cache
| Exact cache | Semantic cache | |
|---|---|---|
| How it matches | Normalized request → hash → cache | Query → embedding → nearest cached query → similarity threshold |
| Correctness | High | Higher risk |
| Hit rate | Lower | Potentially higher |
Use semantic caching only where approximate equivalence is acceptable and has been validated.
7. Conversations and agents
7.1 Conversation cache correctness
Conversational agents add another challenge:
User: What is our refund policy?
Assistant: 30 days.
User: What about enterprise customers?
The answer depends on the previous context, so the query "What about enterprise customers?" alone can't identify the response. You need something like:
conversation_fingerprint = hash(relevant_conversation_state)
key = (
f"conversation-response:"
f"session:{session_id}:"
f"context:{conversation_fingerprint}:"
f"query:{query_hash}"
)7.2 Conversation branching
Imagine a conversation with messages M1 to M4, and the user edits M2. Now you have two branches, and any cached state after the changed message may no longer be valid:
- M1
- M2
- M3
- M4
- M1
- M2 (old)M2 (new)
- M3 (old)M3 (new)
Everything downstream of the changed node must be treated separately.
7.3 Agent state consistency
Agents are even harder:
- Agent state
- Plan
- Tool A
- Tool B
- Final answer
If Tool A changes the state, the cached Tool B result may no longer be valid. So agent caches should include the state version:
key = (
f"tool-result:"
f"tool:{tool_name}:"
f"tool_version:{tool_version}:"
f"state:{state_version}:"
f"input:{input_hash}"
)8. Concurrency
8.1 Lost updates
Multiple workers may modify the same cached state:
| Worker | Reads | Writes |
|---|---|---|
| A | Version 10 | Version 11 |
| B | Version 10 | Version 11 |
Both make changes, both write version 11, and one update is lost. This is a classic race condition.
8.2 Version checks
Use an expected version when updating state:
def update_state(key, expected_version, new_value):
current = get_state(key)
if current.version != expected_version:
raise ConflictError()
write(key, version=expected_version + 1, value=new_value)Now concurrent updates are detected instead of silently overwriting each other.
8.3 Redis WATCH and optimistic transactions
For Redis-backed state, optimistic concurrency can be implemented with transactions and watched keys:
from redis.exceptions import WatchError
with redis.pipeline() as pipe:
while True:
try:
pipe.watch(key)
current = pipe.get(key)
# compute the update from `current`
pipe.multi()
pipe.set(key, new_value)
pipe.execute()
break
except WatchError:
continueThe important idea:
- Read
- Watch
- Prepare update
- Commit only if unchanged
This prevents some lost-update scenarios.
9. Write strategies
9.1 Write-through cache
With cache-aside, the application controls population: it reads from the database and fills the cache. With write-through, the cache layer writes to the source as part of the operation:
- Application
- Database
- Cache
- Application
- Cache
- Database
Conceptually:
def write_user(user):
database.write(user)
cache.set(f"user:{user.id}", user)The exact ordering and failure handling matter.
9.2 The write-through trade-off
| Advantages | Disadvantages |
|---|---|
| The cache stays warm | Every write pays the cache overhead |
| Simpler reads | Failure coordination becomes harder |
Write-through suits frequently read objects and strongly coordinated application state. It's less useful for large AI-generated responses, expensive asynchronous computations and rarely read objects.
9.3 Write-behind
With write-behind, the application writes to the cache, which persists to the database asynchronously:
- Application
- Cache
- Async persistence
- Database
This can improve write performance. But now the cache temporarily holds the newest version of the data, and if it disappears before persistence, the data is lost. So write-behind needs durable queues or other recovery mechanisms when the data matters.
For most AI response caches this isn't necessary, because the cache isn't the source of truth.
9.4 Cache-aside remains a great default
For AI workloads, cache-aside is often the easiest to reason about:
| Operation | Flow |
|---|---|
| Read | Cache → source |
| Write | Source → invalidate or update the cache |
The source stays authoritative. For example:
async def get_document(doc_id):
key = f"document:{doc_id}"
cached = await cache.get(key)
if cached:
return cached
document = await database.get(doc_id)
await cache.set(key, document, ttl=300)
return document10. Event-driven invalidation
10.1 Invalidation through events
For larger systems:
- Database
- Document updated
- Event
- Message broker
- Cache invalidator
The event and its consumer:
event = {
"type": "document.updated",
"document_id": "doc_123",
"version": 42,
}
async def handle(event):
if event["type"] != "document.updated":
return
doc_id = event["document_id"]
await invalidate_document_caches(doc_id)This decouples the write path from cache invalidation.
10.2 Invalidation events should be idempotent
Message brokers can deliver the same event more than once, so document.updated might arrive once, twice or three times. Your handler must process duplicates safely.
A handler that does counter += 1 behaves differently on every duplicate. A handler that calls invalidate(document_id) doesn't, because deleting an already-deleted key is harmless.
Idempotency is critical in distributed invalidation systems.
10.3 Event ordering
Now consider events for document v10 and v11 arriving in the wrong order: v11 first, then v10. If your system processes them blindly, it can move state backwards.
So include versions in events, and ignore anything older than what you've already seen:
event = {"document_id": "doc_123", "version": 11}
if event["version"] <= current_version:
returnThis is a simple but powerful protection against stale events.
11. Generational caching
11.1 Generations
A generation is effectively a namespace version. With generation = 7, keys look like rag:g7:query:abc. When the knowledge changes, the generation becomes 8 and new requests use rag:g8:query:abc.
This makes mass invalidation cheap.
11.2 A global generation
You can keep a single counter:
knowledge_generation = 8
key = f"response:knowledge:{knowledge_generation}:{query_hash}"Incrementing it (knowledge_generation += 1) effectively invalidates every response based on the previous generation.
11.3 Fine-grained vs global invalidation
| Global invalidation | Fine-grained invalidation | |
|---|---|---|
| What happens | Knowledge changes → invalidate everything | Document A changes → invalidate only caches that depend on A |
| Pros | Simple | Efficient |
| Cons | Expensive | Complex |
The right choice depends on dataset size, update frequency, cache volume and correctness requirements.
12. What changed? Invalidation in RAG
12.1 A document changes
Let's walk through a real RAG example. The initial state is document v1, embedding v1 and index v10, with these caches:
embedding:v1:doc123
retrieval:index10:queryABC
response:index10:queryABCNow the document moves to v2. A safe pipeline:
- Document update
- Knowledge version changes
- Re-embed
- Update vector index
- New index version
- New retrieval cache namespace
- New response namespace
The old caches expire naturally.
12.2 Only metadata changes
Suppose the document text stays identical, but its department metadata changes from engineering to finance. The embedding may still be valid, but the retrieval filter may not be. So the embedding cache can stay, while the retrieval cache may need invalidating.
This is why dependency-aware versioning beats blindly invalidating everything.
12.3 Only the prompt changes
Suppose prompt v12 becomes prompt v13. Embeddings, retrieval and reranking may all remain valid, but context → response is affected.
You don't need to rebuild your entire RAG pipeline. You need a new response namespace, and potentially a new prefix cache.
12.4 The model changes
Suppose model A becomes model B. The retrieved evidence may still be valid, but the response cache should not cross model versions:
key = (
f"response:"
f"model:{model_version}:"
f"prompt:{prompt_version}:"
f"context:{context_hash}"
)Now model A and model B naturally have separate caches.
12.5 Summary: what to invalidate
| What changed | Still valid | Needs a new namespace |
|---|---|---|
| Document text | — | Embedding, index, retrieval, response |
| Document metadata only | Embedding | Retrieval, response |
| Prompt | Embedding, retrieval, reranking | Response, possibly prefix cache |
| Model | Retrieved evidence | Response, prefix cache |
12.6 Prompt cache vs response cache
These are fundamentally different:
| Prompt / prefix cache | Response cache | |
|---|---|---|
| Caches | Reusable model computation (system prompt → KV state) | Input → final answer |
| On a model change | Affected | Affected |
Both may be affected by a model change, but the invalidation mechanics are different. Don't treat every AI cache as one generic cache.
13. Memory and tools
13.1 Agent memory invalidation
Agent memory can contain user preferences, past decisions, historical facts, temporary state and long-term memory. Suppose the user says:
"I no longer work at Company A."
The memory user_company = Company A is now stale. Depending on how that memory is used, you should update or invalidate the memory cache, the personalization cache, the conversation context and the agent state.
13.2 Memory versioning
Instead of memory:user123, use a versioned key such as memory:user:user123:version:42. When memory changes, it becomes version 43, and downstream caches can reference memory_version:
key = (
f"agent-response:"
f"user:{user_id}:"
f"memory:{memory_version}:"
f"query:{query_hash}"
)This prevents old personalization from contaminating new responses.
13.3 Tool cache invalidation
Consider get_account_balance(). Caching its result for an hour could be dangerous, because the balance can change at any moment. Use a TTL of a few seconds, or invalidate on every transaction.
The cache policy depends on the tool's semantics.
13.4 Read tools vs write tools
| Tool type | Examples | Caching |
|---|---|---|
| Read | get_weather(), search_products(), get_profile() |
May be safe |
| Write | transfer_money(), send_email(), delete_file(), create_order() |
Caching the result of the action itself can be dangerous; use idempotency keys |
13.5 Idempotency vs caching
These are often confused.
- Caching asks: can I reuse this previous result?
- Idempotency asks: if I execute this operation again, will it safely produce the same effect?
For a payment like charge_customer(), you don't want a cache hit deciding whether the payment happened. You want an idempotency key:
request = {"idempotency_key": "order-123-payment-1"}The backend guarantees that repeating the request doesn't create multiple charges.
14. Security
14.1 Permission changes are invalidation events
Suppose User A has access to document X, and the cache holds response:userA:query123. Then their access is revoked. If you keep serving the cached response, User A may still see information they shouldn't.
So permission changes can be invalidation events:
- ACL change
- Invalidate affected caches
14.2 Authorization should not depend on the cache
Never assume a cache hit means the user is authorized. The authorization layer must stay authoritative. For sensitive data:
- Authenticate
- Authorize
- Determine cache scope
- Cache lookup
A cached result should never bypass access control.
14.3 Cache poisoning
Another dangerous scenario:
- Attacker
- Malicious request
- Cache
- Victims receive the poisoned response
For example, if a cache key doesn't include the tenant, language or authorization scope, one user's result can become another user's result. Cache poisoning can be caused by bad keys, untrusted inputs, incorrect normalization and missing authorization scope.
14.4 Canonicalization
If your application treats "Hello" and " hello " as equivalent, normalize before generating the key:
import hashlib
def normalize_query(query):
return " ".join(query.strip().lower().split())
normalized = normalize_query(query)
query_hash = hashlib.sha256(normalized.encode()).hexdigest()But be careful: normalization must preserve the distinctions that matter. "Apple" and "apple" are not always equivalent.
14.5 Request fingerprints
A production cache often benefits from a structured request fingerprint:
import hashlib
import json
def fingerprint(request):
payload = {
"query": request["query"],
"model": request["model"],
"prompt_version": request["prompt_version"],
"knowledge_version": request["knowledge_version"],
"index_version": request["index_version"],
"tenant": request["tenant"],
}
canonical = json.dumps(payload, sort_keys=True)
return hashlib.sha256(canonical.encode()).hexdigest()Now the same configuration plus the same input gives the same fingerprint, as long as every relevant dependency is represented.
15. Testing correctness
15.1 Cache correctness tests
Don't test only that the cache hits. Test:
- a cache hit with the same input
- a cache miss after a version change
- invalidation after a document update
- tenant isolation
- model-version isolation
- prompt-version isolation
- permission changes
- concurrent updates
- duplicate invalidation events
- out-of-order events
15.2 An example RAG invalidation test
async def test_document_update_invalidates_response():
document_version = 1
index_version = 10
key = build_response_key(
query="refund policy",
document_version=document_version,
index_version=index_version,
)
await cache.set(key, "Old answer")
document_version = 2
index_version = 11
new_key = build_response_key(
query="refund policy",
document_version=document_version,
index_version=index_version,
)
assert await cache.get(new_key) is NoneThe important test isn't "did Redis delete the old key?". It's:
Can the new request accidentally consume an incompatible old result?
15.3 Property-based thinking
For cache correctness, useful invariants include:
- If
model_versionchanges, an old response must never be returned. - If the tenant changes, private cache entries must never cross tenants.
- If
knowledge_versionchanges, old knowledge-dependent responses must not be served. - If authorization is revoked, restricted data must not remain accessible through the cache.
- If a tool's input changes, the previous tool result must not be treated as equivalent.
These are stronger than individual test cases.
16. Measuring staleness
16.1 Observability for correctness
Track cache hits, misses, stale serves, invalidations, invalidation lag, version mismatches, cache conflicts, duplicate events, out-of-order events and authorization cache misses.
Invalidation lag is especially important. It's the time between the data changing and the cache actually becoming invalid:
- Data changes
- Invalidation event
- Cache becomes invalid
If that takes 2 seconds, your system has a 2-second consistency window.
16.2 Measuring cache age
For some systems, store when each entry was created and from which source version:
import time
entry = {
"value": result,
"created_at": time.time(),
"source_version": 42,
}
age = time.time() - entry["created_at"]Now you can monitor P50, P95 and P99 cache age, and the maximum cache age. That is much more informative than the TTL alone.
16.3 Freshness budgets
Instead of saying "TTL = 300", say "maximum acceptable staleness = 30 seconds", and design the architecture around that requirement. The whole path must complete within the budget:
- Source update
- Event
- Consumer
- Invalidate
This becomes a measurable engineering requirement.
16.4 Strong vs eventual consistency in AI
AI systems often don't need strong consistency everywhere. "What are today's trending topics?" is fine with eventual consistency. "What documents am I authorized to access?" is not.
An AI system doesn't have one global consistency requirement. Different components have different requirements.
16.5 A consistency matrix
Create one:
| Component | Staleness tolerance | Strategy |
|---|---|---|
| Public content | High | TTL |
| Tool metadata | High | TTL / version |
| Embeddings | Medium | Versioning |
| Weather | Medium | TTL |
| Vector index | Low | Versioning |
| RAG retrieval | Low | Index version |
| User memory | Low | Event / version |
| Response cache | Depends | Dependency versioning |
| Authorization | Very low | Explicit invalidation |
| Account balance | Very low | No long-lived cache |
This forces the team to make correctness explicit.
17. Invalidation architecture
17.1 A cache invalidation architecture
A mature system can look like:
- Source system
- Change event
- Event bus
- RAG cacheAgent cacheApp cache
- New versionNew versionNew version
- New requests
- Fresh cache
The cache doesn't have to know synchronously about every database mutation. Events propagate the change.
17.2 The AI cache dependency graph
A complete AI system might have:
- Source
- DocumentsUser data
- EmbeddingsMemory
- Vector indexPersonalization
- Retrieval
- Reranking
- Context
- PromptTool state
- LLM
- Response
Each arrow is a dependency. When something changes upstream, every downstream cache needs a strategy.
17.3 Don't invalidate more than necessary
If the prompt changes, you don't necessarily need to invalidate embeddings, the vector index or retrieval; you may only need the response cache and the prefix cache. Similarly, a new embedding model doesn't require invalidating conversation history.
This is why dependency modelling matters.
18. Making cache identity explicit
18.1 A version registry
One simple approach is to keep versions in one place:
VERSIONS = {
"knowledge": 8,
"embedding": 3,
"index": 21,
"reranker": 5,
"prompt": 12,
"model": 7,
"tool_schema": 4,
}
def build_response_key(tenant_id, query_hash, versions):
return (
f"response:"
f"tenant:{tenant_id}:"
f"knowledge:{versions['knowledge']}:"
f"index:{versions['index']}:"
f"prompt:{versions['prompt']}:"
f"model:{versions['model']}:"
f"query:{query_hash}"
)This makes cache identity explicit.
18.2 Better: version objects
For larger systems, use a typed object:
from dataclasses import dataclass
@dataclass(frozen=True)
class CacheVersions:
knowledge: int
embedding: int
index: int
reranker: int
prompt: int
model: int
versions = CacheVersions(
knowledge=8,
embedding=3,
index=21,
reranker=5,
prompt=12,
model=7,
)Now cache builders receive an explicit dependency object, which is easier to test.
18.3 A production cache key builder
class CacheKeyBuilder:
def response_key(self, tenant_id, query_hash, versions):
return (
"response:"
f"tenant={tenant_id}:"
f"knowledge={versions.knowledge}:"
f"index={versions.index}:"
f"prompt={versions.prompt}:"
f"model={versions.model}:"
f"query={query_hash}"
)Now every service uses the same rules. This avoids the dangerous situation where service A keys on model + prompt, service B keys on model + prompt + index, and nobody notices.
18.4 A cache contract
A strong production team defines a contract for every cache: name, purpose, owner, source of truth, key format, dependencies, TTL, invalidation trigger, consistency requirement, maximum size, security scope, failure behaviour and observability. For example:
| Field | RAG response cache |
|---|---|
| Source | LLM + retrieved knowledge |
| Dependencies | Model, prompt, knowledge, index, tenant, query |
| TTL | 5 minutes |
| Invalidation | Knowledge, index, prompt or model version change |
| Failure | Fall through to generation |
| Security | Tenant-scoped |
This turns caching from tribal knowledge into architecture.
19. Before you ship
19.1 The cache correctness checklist
Before shipping an AI cache, ask:
| Area | Question |
|---|---|
| Identity | What makes two requests equivalent? |
| Dependencies | What data affects the result? |
| Versioning | Which versions affect correctness? |
| Freshness | How stale can this safely be? |
| Invalidation | What event makes this invalid? |
| Security | Can this value cross tenants or users? |
| Concurrency | What if two workers update it? |
| Failure | What happens if Redis disappears? |
| Stampede | What happens if 100K requests miss at once? |
| Observability | Can we detect stale or incorrect cache behaviour? |
19.2 The biggest mistakes
- Using a TTL as the only invalidation mechanism.
- Using
hash(query)as the entire AI response cache key. - Ignoring the model, prompt, knowledge and index versions.
- Sharing cache entries across tenants.
- Caching authorization decisions too aggressively.
- Assuming semantic similarity means semantic equivalence.
- Ignoring concurrent writes.
- Invalidating everything for every change.
- Having no invalidation observability.
- Treating cached AI output as inherently trustworthy.
20. The full picture
20.1 The production principle
Here's the mindset to remember:
Every cached result should have a clearly defined validity boundary.
That boundary can be a TTL, a version, an event, a dependency, a transaction, a session, a user, a tenant, a model or a knowledge snapshot.
If you can't explain what makes a cached result invalid, you probably shouldn't ship that cache yet.
20.2 The complete correctness model
A production AI cache entry has three parts:
- Cache entry
- Valuethe resultIdentitytenant, user, query, model, dependenciesFreshnessTTL, version, event
A cache hit is safe only when all four of these hold:
- The identity matches.
- The dependencies match.
- The freshness policy allows serving.
- Authorization still allows serving.
That's the real definition of a safe cache hit.
20.3 Part 8 takeaways
If you remember only these:
- TTL is not invalidation. A TTL controls lifetime; it doesn't know when your data changed.
- Cache keys are correctness boundaries. A missing dependency can produce an incorrect cache hit.
- Versioning is extremely powerful for AI. Version models, prompts, embeddings, indexes, knowledge, tools and schemas.
- AI caches form dependency graphs. Document → embedding → retrieval → context → prompt → response: changes propagate downstream.
- Not every change requires invalidating everything. Understand the graph, and invalidate only what depends on the changed component.
- Authorization is not caching. Never let a cache hit bypass permission checks.
- Semantic caches need extra caution. Similar doesn't necessarily mean equivalent.
- Distributed invalidation must handle duplicates and ordering. Events can arrive twice, late or out of order. Design for it.
- Measure staleness. Don't monitor only the hit rate; monitor cache age, invalidation lag, version mismatches and stale serves.
- Correctness comes before hit rate. A 99% hit rate is worthless if 10% of those hits are wrong.
20.4 The core principle
Caching isn't fundamentally about storing data. It's about making this statement true:
"This result is equivalent to what I would compute now."
If you can prove that, a cache hit means a correct result, and you have a useful cache. If you can't, you're gambling with correctness. And in AI systems, that gamble compounds:
- Wrong knowledge
- Wrong retrieval
- Wrong context
- Wrong reasoning
- Wrong answer
So the production rule is simple:
Never optimize cache hit rate at the expense of cache correctness.
What's next?
Part 9: AI cache observability, evaluation and cost optimization
We've now answered what to cache, how to cache it, how to scale it, how to invalidate it and how to keep it correct. One question is left:
Is the cache actually helping?
Part 9 focuses on the measurement layer:
- hit rate, miss rate and hit quality
- staleness, latency and TTFT
- tokens saved and GPU compute saved
- LLM, embedding and tool calls avoided
- cost saved, and cache memory and eviction efficiency
- hot-key and stampede detection
- cache effectiveness and ROI
And most importantly: how do you prove that your caching architecture is making the AI system faster, cheaper and better?
Because in production, a cache that exists is not the same as a cache that creates value. The real goal:
- Cache
- Lower latencyLower costLower compute
- Same or better quality
That is where caching becomes a real AI engineering optimization, rather than just another infrastructure component.
Next: Part 9, Cache observability and cost covers measuring latency, tokens, calls avoided, cost, ROI and hit quality.