HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 8: Cache Invalidation and Correctness

Fast is useless if the answer is wrong: TTLs, versioned keys, dependency graphs, events, concurrency and testing for correct AI caches.

Haribaskar Dhanabalan19 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

Fast is useless if the answer is wrong.

This is Part 8 of the AI Caching Playbook. The earlier parts:

Caching sounds simple:

  1. Request
  2. Cache
  3. Hit
  4. Return result

But production systems aren't static. Data changes. Documents get updated. Users change permissions. Models are upgraded. Prompts change. Embeddings are regenerated. Vector indexes are rebuilt. Tools return different results. Agents create new state.

And suddenly the cache holds an answer that was correct five minutes ago. That is the fundamental caching problem:

How do you reuse expensive computation without reusing information that is no longer valid?

This matters especially in AI systems. Consider a RAG application:

  1. Document v1
  2. Embedding v1
  3. Vector index v1
  4. Retrieval
  5. Reranking
  6. Context
  7. LLM
  8. Answer

Now the document changes to v2. What happens to the embedding, the vector index, the retrieval cache, the reranking cache, the context cache and the response cache?

If you don't answer that question explicitly, your cache architecture is incomplete.

1. The cache correctness problem

1.1 Stale data

Suppose your application caches user:123 → Premium. Then the user downgrades from Premium to Free, but the cache still says Premium. The application may keep treating them as a premium user.

That's stale data. Now imagine the same problem with AI:

  1. Knowledge base
  2. Old information
  3. Cached answer
  4. User receives an outdated answer

The model isn't necessarily wrong. The cache is.

1.2 Cache invalidation

Cache invalidation means:

Making cached data unusable when its underlying dependencies change.

There are several major approaches: TTLs, explicit deletion, versioning, event-driven invalidation, dependency-based invalidation, namespace invalidation, generational invalidation, write-through and write-behind.

No single strategy works for everything.

2. TTL

2.1 Time to live

The simplest strategy:

redis.set(key, value, ex=300)

The entry expires after 300 seconds, which gives you maximum staleness ≈ TTL.

But there's an important catch: a TTL doesn't know whether your data changed. With a 1-hour TTL, if the document changes 10 seconds after caching, the cache can stay stale for almost 59 minutes and 50 seconds.

TTL controls lifetime, not correctness.

2.2 When TTL is enough

A TTL works well when stale data is acceptable: weather, popular content, recommendations, analytics, non-critical metadata, public dashboards. For example:

redis.set("trending_topics", data, ex=60)

A one-minute-old trending list may be perfectly fine.

2.3 When TTL is not enough

Be careful with authorization, permissions, financial state, inventory, private data, compliance information, security configuration and critical business rules. A five-minute-stale value might be unacceptable.

For these systems, a TTL should not be your only correctness mechanism.

3. Explicit and dependency-aware invalidation

3.1 Explicit invalidation

The next strategy is direct deletion. When document:123 changes:

redis.delete("document:123")

Simple. But AI systems have dependencies. That document may have produced embedding:123, retrieval:abc, rerank:xyz, context:qwe and response:123. Deleting only document:123 doesn't invalidate the derived results.

3.2 Dependency graphs

Think of cached computation as a graph:

  1. Document
  2. Embedding
  3. Vector index
  4. Retrieval
  5. Reranking
  6. Context
  7. Prompt
  8. Response

If the document changes, everything downstream may become invalid. This is a dependency problem.

3.3 Dependency-aware invalidation

You can track dependencies explicitly:

dependencies = {
    "response:abc": ["context:xyz", "model:v5", "prompt:v12"],
    "context:xyz": ["retrieval:qwe"],
    "retrieval:qwe": ["document:123", "document:456"],
}
 
 
def invalidate_document(document_id):
    affected = find_dependents(document_id)
 
    for key in affected:
        redis.delete(key)

This gives you precise invalidation. The problem: at large scale, maintaining dependency graphs can get expensive.

4. Version-based invalidation

4.1 Change the version, not the keys

This is one of the most useful patterns for AI. Instead of deleting 10 million cache entries, change the version: knowledge:v1 becomes knowledge:v2, and the version is part of every key:

key = (
    f"rag:"
    f"knowledge:{knowledge_version}:"
    f"query:{query_hash}"
)

Old entries look like rag:knowledge:1:abc and new ones like rag:knowledge:2:abc. The old cache simply expires naturally.

4.2 Why versioning works so well for AI

AI systems have many versioned dependencies: the model, prompt, embedding model, index, reranker, tools, schema and knowledge base. Instead of trying to understand every invalidation relationship:

  1. A dependency changes
  2. Its version changes
  3. New cache namespace

This is much easier to reason about.

4.3 Embedding versioning

Suppose you start with text-embedding-model-v1 and migrate to text-embedding-model-v2. An embedding produced by v1 must never be mistaken for a v2 embedding. A good key:

key = (
    f"embedding:"
    f"model:{model_name}:"
    f"version:{model_version}:"
    f"preprocess:{preprocess_version}:"
    f"text:{text_hash}"
)

Now changing the model automatically creates a new cache namespace.

4.4 Retrieval versioning

Suppose your vector index moves from index:v10 to index:v11. The retrieval cache should include the index version:

key = (
    f"retrieval:"
    f"index:{index_version}:"
    f"query:{query_hash}:"
    f"topk:{top_k}"
)

Without index versioning:

  1. New index
  2. Old retrieval result
  3. Incorrect evidence

That's a subtle but serious RAG bug.

4.5 Response cache versioning

A final AI response can depend on many things:

key = (
    f"response:"
    f"tenant:{tenant_id}:"
    f"model:{model_version}:"
    f"prompt:{prompt_version}:"
    f"knowledge:{knowledge_version}:"
    f"index:{index_version}:"
    f"query:{query_hash}"
)

Now a change to the model, prompt, knowledge or index automatically creates a new response cache.

5. The cache key is a correctness contract

5.1 What a key really says

A cache key isn't just a lookup identifier. It defines:

Under what conditions is this cached result equivalent to another request?

response:abc123 doesn't tell us enough. This does:

response:tenant:acme:model:v5:prompt:v12:knowledge:v8:index:v21:query:abc123

Every component answers a correctness question.

5.2 What should go into an AI cache key?

Depending on the cache:

Area Possible key components
Scope Tenant, user scope
Input Query, normalized query, filters, top-K
Model Model, model version, temperature and other relevant generation settings
Prompt Prompt version
Retrieval Embedding model and version, preprocessing version, index version, reranker version, knowledge version
Tools Tool version, schema version

Don't blindly include everything. Only include values that actually affect the cached output.

5.3 Cache-key explosion

There is a downside. If you include temperature, top-K, model, prompt, tenant, user, region and language, the same question can create hundreds of different cache keys, and your hit rate collapses.

So the goal isn't to put everything into the key. The goal is to put every correctness-relevant dependency into the key.

6. Semantic equivalence

6.1 Similar or the same?

"What is the refund policy?" and "Can you tell me your refund policy?" are semantically similar. Should they share a cache entry?

For an exact cache: no. For a semantic cache: potentially. But semantic caching introduces another correctness problem: two questions can be similar without being equivalent.

6.2 Semantic cache risk

"What is the price of Product A?" and "What was the price of Product A last year?" may have very high embedding similarity, but the answers are completely different.

So semantic similarity ≠ semantic equivalence. This is why semantic response caches need conservative thresholds and often extra validation.

6.3 Exact cache vs semantic cache

Exact cache Semantic cache
How it matches Normalized request → hash → cache Query → embedding → nearest cached query → similarity threshold
Correctness High Higher risk
Hit rate Lower Potentially higher

Use semantic caching only where approximate equivalence is acceptable and has been validated.

7. Conversations and agents

7.1 Conversation cache correctness

Conversational agents add another challenge:

User: What is our refund policy?

Assistant: 30 days.

User: What about enterprise customers?

The answer depends on the previous context, so the query "What about enterprise customers?" alone can't identify the response. You need something like:

conversation_fingerprint = hash(relevant_conversation_state)
 
key = (
    f"conversation-response:"
    f"session:{session_id}:"
    f"context:{conversation_fingerprint}:"
    f"query:{query_hash}"
)

7.2 Conversation branching

Imagine a conversation with messages M1 to M4, and the user edits M2. Now you have two branches, and any cached state after the changed message may no longer be valid:

  1. M1
  2. M2
  3. M3
  4. M4
Original
  1. M1
  2. M2 (old)M2 (new)
  3. M3 (old)M3 (new)
After editing M2

Everything downstream of the changed node must be treated separately.

7.3 Agent state consistency

Agents are even harder:

  1. Agent state
  2. Plan
  3. Tool A
  4. Tool B
  5. Final answer

If Tool A changes the state, the cached Tool B result may no longer be valid. So agent caches should include the state version:

key = (
    f"tool-result:"
    f"tool:{tool_name}:"
    f"tool_version:{tool_version}:"
    f"state:{state_version}:"
    f"input:{input_hash}"
)

8. Concurrency

8.1 Lost updates

Multiple workers may modify the same cached state:

Worker Reads Writes
A Version 10 Version 11
B Version 10 Version 11

Both make changes, both write version 11, and one update is lost. This is a classic race condition.

8.2 Version checks

Use an expected version when updating state:

def update_state(key, expected_version, new_value):
    current = get_state(key)
 
    if current.version != expected_version:
        raise ConflictError()
 
    write(key, version=expected_version + 1, value=new_value)

Now concurrent updates are detected instead of silently overwriting each other.

8.3 Redis WATCH and optimistic transactions

For Redis-backed state, optimistic concurrency can be implemented with transactions and watched keys:

from redis.exceptions import WatchError
 
with redis.pipeline() as pipe:
    while True:
        try:
            pipe.watch(key)
            current = pipe.get(key)
 
            # compute the update from `current`
 
            pipe.multi()
            pipe.set(key, new_value)
            pipe.execute()
            break
        except WatchError:
            continue

The important idea:

  1. Read
  2. Watch
  3. Prepare update
  4. Commit only if unchanged

This prevents some lost-update scenarios.

9. Write strategies

9.1 Write-through cache

With cache-aside, the application controls population: it reads from the database and fills the cache. With write-through, the cache layer writes to the source as part of the operation:

  1. Application
  2. Database
  3. Cache
Cache-aside
  1. Application
  2. Cache
  3. Database
Write-through

Conceptually:

def write_user(user):
    database.write(user)
    cache.set(f"user:{user.id}", user)

The exact ordering and failure handling matter.

9.2 The write-through trade-off

Advantages Disadvantages
The cache stays warm Every write pays the cache overhead
Simpler reads Failure coordination becomes harder

Write-through suits frequently read objects and strongly coordinated application state. It's less useful for large AI-generated responses, expensive asynchronous computations and rarely read objects.

9.3 Write-behind

With write-behind, the application writes to the cache, which persists to the database asynchronously:

  1. Application
  2. Cache
  3. Async persistence
  4. Database

This can improve write performance. But now the cache temporarily holds the newest version of the data, and if it disappears before persistence, the data is lost. So write-behind needs durable queues or other recovery mechanisms when the data matters.

For most AI response caches this isn't necessary, because the cache isn't the source of truth.

9.4 Cache-aside remains a great default

For AI workloads, cache-aside is often the easiest to reason about:

Operation Flow
Read Cache → source
Write Source → invalidate or update the cache

The source stays authoritative. For example:

async def get_document(doc_id):
    key = f"document:{doc_id}"
 
    cached = await cache.get(key)
    if cached:
        return cached
 
    document = await database.get(doc_id)
    await cache.set(key, document, ttl=300)
    return document

10. Event-driven invalidation

10.1 Invalidation through events

For larger systems:

  1. Database
  2. Document updated
  3. Event
  4. Message broker
  5. Cache invalidator

The event and its consumer:

event = {
    "type": "document.updated",
    "document_id": "doc_123",
    "version": 42,
}
 
 
async def handle(event):
    if event["type"] != "document.updated":
        return
 
    doc_id = event["document_id"]
    await invalidate_document_caches(doc_id)

This decouples the write path from cache invalidation.

10.2 Invalidation events should be idempotent

Message brokers can deliver the same event more than once, so document.updated might arrive once, twice or three times. Your handler must process duplicates safely.

A handler that does counter += 1 behaves differently on every duplicate. A handler that calls invalidate(document_id) doesn't, because deleting an already-deleted key is harmless.

Idempotency is critical in distributed invalidation systems.

10.3 Event ordering

Now consider events for document v10 and v11 arriving in the wrong order: v11 first, then v10. If your system processes them blindly, it can move state backwards.

So include versions in events, and ignore anything older than what you've already seen:

event = {"document_id": "doc_123", "version": 11}
 
if event["version"] <= current_version:
    return

This is a simple but powerful protection against stale events.

11. Generational caching

11.1 Generations

A generation is effectively a namespace version. With generation = 7, keys look like rag:g7:query:abc. When the knowledge changes, the generation becomes 8 and new requests use rag:g8:query:abc.

This makes mass invalidation cheap.

11.2 A global generation

You can keep a single counter:

knowledge_generation = 8
 
key = f"response:knowledge:{knowledge_generation}:{query_hash}"

Incrementing it (knowledge_generation += 1) effectively invalidates every response based on the previous generation.

11.3 Fine-grained vs global invalidation

Global invalidation Fine-grained invalidation
What happens Knowledge changes → invalidate everything Document A changes → invalidate only caches that depend on A
Pros Simple Efficient
Cons Expensive Complex

The right choice depends on dataset size, update frequency, cache volume and correctness requirements.

12. What changed? Invalidation in RAG

12.1 A document changes

Let's walk through a real RAG example. The initial state is document v1, embedding v1 and index v10, with these caches:

embedding:v1:doc123
retrieval:index10:queryABC
response:index10:queryABC

Now the document moves to v2. A safe pipeline:

  1. Document update
  2. Knowledge version changes
  3. Re-embed
  4. Update vector index
  5. New index version
  6. New retrieval cache namespace
  7. New response namespace

The old caches expire naturally.

12.2 Only metadata changes

Suppose the document text stays identical, but its department metadata changes from engineering to finance. The embedding may still be valid, but the retrieval filter may not be. So the embedding cache can stay, while the retrieval cache may need invalidating.

This is why dependency-aware versioning beats blindly invalidating everything.

12.3 Only the prompt changes

Suppose prompt v12 becomes prompt v13. Embeddings, retrieval and reranking may all remain valid, but context → response is affected.

You don't need to rebuild your entire RAG pipeline. You need a new response namespace, and potentially a new prefix cache.

12.4 The model changes

Suppose model A becomes model B. The retrieved evidence may still be valid, but the response cache should not cross model versions:

key = (
    f"response:"
    f"model:{model_version}:"
    f"prompt:{prompt_version}:"
    f"context:{context_hash}"
)

Now model A and model B naturally have separate caches.

12.5 Summary: what to invalidate

What changed Still valid Needs a new namespace
Document text — Embedding, index, retrieval, response
Document metadata only Embedding Retrieval, response
Prompt Embedding, retrieval, reranking Response, possibly prefix cache
Model Retrieved evidence Response, prefix cache

12.6 Prompt cache vs response cache

These are fundamentally different:

Prompt / prefix cache Response cache
Caches Reusable model computation (system prompt → KV state) Input → final answer
On a model change Affected Affected

Both may be affected by a model change, but the invalidation mechanics are different. Don't treat every AI cache as one generic cache.

13. Memory and tools

13.1 Agent memory invalidation

Agent memory can contain user preferences, past decisions, historical facts, temporary state and long-term memory. Suppose the user says:

"I no longer work at Company A."

The memory user_company = Company A is now stale. Depending on how that memory is used, you should update or invalidate the memory cache, the personalization cache, the conversation context and the agent state.

13.2 Memory versioning

Instead of memory:user123, use a versioned key such as memory:user:user123:version:42. When memory changes, it becomes version 43, and downstream caches can reference memory_version:

key = (
    f"agent-response:"
    f"user:{user_id}:"
    f"memory:{memory_version}:"
    f"query:{query_hash}"
)

This prevents old personalization from contaminating new responses.

13.3 Tool cache invalidation

Consider get_account_balance(). Caching its result for an hour could be dangerous, because the balance can change at any moment. Use a TTL of a few seconds, or invalidate on every transaction.

The cache policy depends on the tool's semantics.

13.4 Read tools vs write tools

Tool type Examples Caching
Read get_weather(), search_products(), get_profile() May be safe
Write transfer_money(), send_email(), delete_file(), create_order() Caching the result of the action itself can be dangerous; use idempotency keys

13.5 Idempotency vs caching

These are often confused.

  • Caching asks: can I reuse this previous result?
  • Idempotency asks: if I execute this operation again, will it safely produce the same effect?

For a payment like charge_customer(), you don't want a cache hit deciding whether the payment happened. You want an idempotency key:

request = {"idempotency_key": "order-123-payment-1"}

The backend guarantees that repeating the request doesn't create multiple charges.

14. Security

14.1 Permission changes are invalidation events

Suppose User A has access to document X, and the cache holds response:userA:query123. Then their access is revoked. If you keep serving the cached response, User A may still see information they shouldn't.

So permission changes can be invalidation events:

  1. ACL change
  2. Invalidate affected caches

14.2 Authorization should not depend on the cache

Never assume a cache hit means the user is authorized. The authorization layer must stay authoritative. For sensitive data:

  1. Authenticate
  2. Authorize
  3. Determine cache scope
  4. Cache lookup

A cached result should never bypass access control.

14.3 Cache poisoning

Another dangerous scenario:

  1. Attacker
  2. Malicious request
  3. Cache
  4. Victims receive the poisoned response

For example, if a cache key doesn't include the tenant, language or authorization scope, one user's result can become another user's result. Cache poisoning can be caused by bad keys, untrusted inputs, incorrect normalization and missing authorization scope.

14.4 Canonicalization

If your application treats "Hello" and " hello " as equivalent, normalize before generating the key:

import hashlib
 
 
def normalize_query(query):
    return " ".join(query.strip().lower().split())
 
 
normalized = normalize_query(query)
query_hash = hashlib.sha256(normalized.encode()).hexdigest()

But be careful: normalization must preserve the distinctions that matter. "Apple" and "apple" are not always equivalent.

14.5 Request fingerprints

A production cache often benefits from a structured request fingerprint:

import hashlib
import json
 
 
def fingerprint(request):
    payload = {
        "query": request["query"],
        "model": request["model"],
        "prompt_version": request["prompt_version"],
        "knowledge_version": request["knowledge_version"],
        "index_version": request["index_version"],
        "tenant": request["tenant"],
    }
    canonical = json.dumps(payload, sort_keys=True)
    return hashlib.sha256(canonical.encode()).hexdigest()

Now the same configuration plus the same input gives the same fingerprint, as long as every relevant dependency is represented.

15. Testing correctness

15.1 Cache correctness tests

Don't test only that the cache hits. Test:

  • a cache hit with the same input
  • a cache miss after a version change
  • invalidation after a document update
  • tenant isolation
  • model-version isolation
  • prompt-version isolation
  • permission changes
  • concurrent updates
  • duplicate invalidation events
  • out-of-order events

15.2 An example RAG invalidation test

async def test_document_update_invalidates_response():
    document_version = 1
    index_version = 10
 
    key = build_response_key(
        query="refund policy",
        document_version=document_version,
        index_version=index_version,
    )
    await cache.set(key, "Old answer")
 
    document_version = 2
    index_version = 11
 
    new_key = build_response_key(
        query="refund policy",
        document_version=document_version,
        index_version=index_version,
    )
 
    assert await cache.get(new_key) is None

The important test isn't "did Redis delete the old key?". It's:

Can the new request accidentally consume an incompatible old result?

15.3 Property-based thinking

For cache correctness, useful invariants include:

  • If model_version changes, an old response must never be returned.
  • If the tenant changes, private cache entries must never cross tenants.
  • If knowledge_version changes, old knowledge-dependent responses must not be served.
  • If authorization is revoked, restricted data must not remain accessible through the cache.
  • If a tool's input changes, the previous tool result must not be treated as equivalent.

These are stronger than individual test cases.

16. Measuring staleness

16.1 Observability for correctness

Track cache hits, misses, stale serves, invalidations, invalidation lag, version mismatches, cache conflicts, duplicate events, out-of-order events and authorization cache misses.

Invalidation lag is especially important. It's the time between the data changing and the cache actually becoming invalid:

  1. Data changes
  2. Invalidation event
  3. Cache becomes invalid

If that takes 2 seconds, your system has a 2-second consistency window.

16.2 Measuring cache age

For some systems, store when each entry was created and from which source version:

import time
 
entry = {
    "value": result,
    "created_at": time.time(),
    "source_version": 42,
}
 
age = time.time() - entry["created_at"]

Now you can monitor P50, P95 and P99 cache age, and the maximum cache age. That is much more informative than the TTL alone.

16.3 Freshness budgets

Instead of saying "TTL = 300", say "maximum acceptable staleness = 30 seconds", and design the architecture around that requirement. The whole path must complete within the budget:

  1. Source update
  2. Event
  3. Consumer
  4. Invalidate
Must complete in under 30 seconds

This becomes a measurable engineering requirement.

16.4 Strong vs eventual consistency in AI

AI systems often don't need strong consistency everywhere. "What are today's trending topics?" is fine with eventual consistency. "What documents am I authorized to access?" is not.

An AI system doesn't have one global consistency requirement. Different components have different requirements.

16.5 A consistency matrix

Create one:

Component Staleness tolerance Strategy
Public content High TTL
Tool metadata High TTL / version
Embeddings Medium Versioning
Weather Medium TTL
Vector index Low Versioning
RAG retrieval Low Index version
User memory Low Event / version
Response cache Depends Dependency versioning
Authorization Very low Explicit invalidation
Account balance Very low No long-lived cache

This forces the team to make correctness explicit.

17. Invalidation architecture

17.1 A cache invalidation architecture

A mature system can look like:

  1. Source system
  2. Change event
  3. Event bus
  4. RAG cacheAgent cacheApp cache
  5. New versionNew versionNew version
  6. New requests
  7. Fresh cache

The cache doesn't have to know synchronously about every database mutation. Events propagate the change.

17.2 The AI cache dependency graph

A complete AI system might have:

  1. Source
  2. DocumentsUser data
  3. EmbeddingsMemory
  4. Vector indexPersonalization
  5. Retrieval
  6. Reranking
  7. Context
  8. PromptTool state
  9. LLM
  10. Response

Each arrow is a dependency. When something changes upstream, every downstream cache needs a strategy.

17.3 Don't invalidate more than necessary

If the prompt changes, you don't necessarily need to invalidate embeddings, the vector index or retrieval; you may only need the response cache and the prefix cache. Similarly, a new embedding model doesn't require invalidating conversation history.

This is why dependency modelling matters.

18. Making cache identity explicit

18.1 A version registry

One simple approach is to keep versions in one place:

VERSIONS = {
    "knowledge": 8,
    "embedding": 3,
    "index": 21,
    "reranker": 5,
    "prompt": 12,
    "model": 7,
    "tool_schema": 4,
}
 
 
def build_response_key(tenant_id, query_hash, versions):
    return (
        f"response:"
        f"tenant:{tenant_id}:"
        f"knowledge:{versions['knowledge']}:"
        f"index:{versions['index']}:"
        f"prompt:{versions['prompt']}:"
        f"model:{versions['model']}:"
        f"query:{query_hash}"
    )

This makes cache identity explicit.

18.2 Better: version objects

For larger systems, use a typed object:

from dataclasses import dataclass
 
 
@dataclass(frozen=True)
class CacheVersions:
    knowledge: int
    embedding: int
    index: int
    reranker: int
    prompt: int
    model: int
 
 
versions = CacheVersions(
    knowledge=8,
    embedding=3,
    index=21,
    reranker=5,
    prompt=12,
    model=7,
)

Now cache builders receive an explicit dependency object, which is easier to test.

18.3 A production cache key builder

class CacheKeyBuilder:
    def response_key(self, tenant_id, query_hash, versions):
        return (
            "response:"
            f"tenant={tenant_id}:"
            f"knowledge={versions.knowledge}:"
            f"index={versions.index}:"
            f"prompt={versions.prompt}:"
            f"model={versions.model}:"
            f"query={query_hash}"
        )

Now every service uses the same rules. This avoids the dangerous situation where service A keys on model + prompt, service B keys on model + prompt + index, and nobody notices.

18.4 A cache contract

A strong production team defines a contract for every cache: name, purpose, owner, source of truth, key format, dependencies, TTL, invalidation trigger, consistency requirement, maximum size, security scope, failure behaviour and observability. For example:

Field RAG response cache
Source LLM + retrieved knowledge
Dependencies Model, prompt, knowledge, index, tenant, query
TTL 5 minutes
Invalidation Knowledge, index, prompt or model version change
Failure Fall through to generation
Security Tenant-scoped

This turns caching from tribal knowledge into architecture.

19. Before you ship

19.1 The cache correctness checklist

Before shipping an AI cache, ask:

Area Question
Identity What makes two requests equivalent?
Dependencies What data affects the result?
Versioning Which versions affect correctness?
Freshness How stale can this safely be?
Invalidation What event makes this invalid?
Security Can this value cross tenants or users?
Concurrency What if two workers update it?
Failure What happens if Redis disappears?
Stampede What happens if 100K requests miss at once?
Observability Can we detect stale or incorrect cache behaviour?

19.2 The biggest mistakes

  1. Using a TTL as the only invalidation mechanism.
  2. Using hash(query) as the entire AI response cache key.
  3. Ignoring the model, prompt, knowledge and index versions.
  4. Sharing cache entries across tenants.
  5. Caching authorization decisions too aggressively.
  6. Assuming semantic similarity means semantic equivalence.
  7. Ignoring concurrent writes.
  8. Invalidating everything for every change.
  9. Having no invalidation observability.
  10. Treating cached AI output as inherently trustworthy.

20. The full picture

20.1 The production principle

Here's the mindset to remember:

Every cached result should have a clearly defined validity boundary.

That boundary can be a TTL, a version, an event, a dependency, a transaction, a session, a user, a tenant, a model or a knowledge snapshot.

If you can't explain what makes a cached result invalid, you probably shouldn't ship that cache yet.

20.2 The complete correctness model

A production AI cache entry has three parts:

  1. Cache entry
  2. Valuethe resultIdentitytenant, user, query, model, dependenciesFreshnessTTL, version, event

A cache hit is safe only when all four of these hold:

  1. The identity matches.
  2. The dependencies match.
  3. The freshness policy allows serving.
  4. Authorization still allows serving.

That's the real definition of a safe cache hit.

20.3 Part 8 takeaways

If you remember only these:

  1. TTL is not invalidation. A TTL controls lifetime; it doesn't know when your data changed.
  2. Cache keys are correctness boundaries. A missing dependency can produce an incorrect cache hit.
  3. Versioning is extremely powerful for AI. Version models, prompts, embeddings, indexes, knowledge, tools and schemas.
  4. AI caches form dependency graphs. Document → embedding → retrieval → context → prompt → response: changes propagate downstream.
  5. Not every change requires invalidating everything. Understand the graph, and invalidate only what depends on the changed component.
  6. Authorization is not caching. Never let a cache hit bypass permission checks.
  7. Semantic caches need extra caution. Similar doesn't necessarily mean equivalent.
  8. Distributed invalidation must handle duplicates and ordering. Events can arrive twice, late or out of order. Design for it.
  9. Measure staleness. Don't monitor only the hit rate; monitor cache age, invalidation lag, version mismatches and stale serves.
  10. Correctness comes before hit rate. A 99% hit rate is worthless if 10% of those hits are wrong.

20.4 The core principle

Caching isn't fundamentally about storing data. It's about making this statement true:

"This result is equivalent to what I would compute now."

If you can prove that, a cache hit means a correct result, and you have a useful cache. If you can't, you're gambling with correctness. And in AI systems, that gamble compounds:

  1. Wrong knowledge
  2. Wrong retrieval
  3. Wrong context
  4. Wrong reasoning
  5. Wrong answer

So the production rule is simple:

Never optimize cache hit rate at the expense of cache correctness.

What's next?

Part 9: AI cache observability, evaluation and cost optimization

We've now answered what to cache, how to cache it, how to scale it, how to invalidate it and how to keep it correct. One question is left:

Is the cache actually helping?

Part 9 focuses on the measurement layer:

  • hit rate, miss rate and hit quality
  • staleness, latency and TTFT
  • tokens saved and GPU compute saved
  • LLM, embedding and tool calls avoided
  • cost saved, and cache memory and eviction efficiency
  • hot-key and stampede detection
  • cache effectiveness and ROI

And most importantly: how do you prove that your caching architecture is making the AI system faster, cheaper and better?

Because in production, a cache that exists is not the same as a cache that creates value. The real goal:

  1. Cache
  2. Lower latencyLower costLower compute
  3. Same or better quality

That is where caching becomes a real AI engineering optimization, rather than just another infrastructure component.

Next: Part 9, Cache observability and cost covers measuring latency, tokens, calls avoided, cost, ROI and hit quality.