HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 7: Distributed AI Caching

Scaling AI caches: Redis Cluster sharding, hot keys, L1/L2 caches, stampedes, invalidation at scale, multi-region, failure handling and cost.

Haribaskar Dhanabalan25 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

What happens when your cache becomes a distributed system?

This is Part 7 of the AI Caching Playbook. The earlier parts built caching up from the ground:

But eventually you hit another problem. Your AI system grows, and one Redis instance is no longer enough:

  1. Millions of users
  2. Millions of requests
  3. Billions of cache operations
  4. Huge cache dataset
  5. Multiple regions
  6. Multiple Redis nodes

Now caching is no longer just an optimization. It becomes a distributed-systems problem. You need to think about sharding, replication, failover, consistency, invalidation, hot keys, cache stampedes, distributed locks, multi-region routing, memory limits, eviction, observability and disaster recovery.

This is where many AI systems get caching wrong.

1. From one node to a cluster

1.1 The single-node cache problem

A simple architecture looks like:

  1. Application
  2. Redis
  3. Database

This works beautifully at small scale. But eventually that one Redis instance hits its CPU, memory and network limits, and it is a single availability risk. One machine is handling everything.

If that Redis instance goes down, the database suddenly receives all the traffic:

  1. Application
  2. Redis ✗
  3. Database

That can create a dangerous cascade:

  1. Cache failure
  2. Cache misses
  3. Database traffic increases
  4. Database slows
  5. Application slows
  6. Requests retry
  7. More traffic
  8. System failure

This is one of the most important ideas in production caching:

A cache failure should not become an application failure.

1.2 Distributed cache architecture

A larger system might look like:

  1. Application
  2. Cache client
  3. Redis-1Redis-2Redis-3
  4. Database

Now the cache data is distributed. In Redis Cluster, keys are spread across 16,384 hash slots, and each cluster node is responsible for a subset of those slots:

  1. Hash(key)
  2. Hash slot
  3. Redis node

1.3 Sharding

Sharding means splitting the cache dataset across multiple nodes. Instead of one Redis instance holding 1 billion keys, a three-node cluster holds roughly a third each:

  1. Redis Cluster
  2. Node A~⅓ of keysNode B~⅓ of keysNode C~⅓ of keys

The application doesn't need to know which node owns a key, as long as the client understands the cluster topology. The logical operation stays redis.get(key), and the cluster-aware client routes it to the correct shard.

Redis Cluster uses hash slots rather than classic consistent hashing.

1.4 How key distribution works

Conceptually:

  1. Key
  2. CRC16
  3. mod 16384
  4. Hash slot
  5. Node

For example (slot numbers are illustrative):

Key Hash slot Node
user:123 8192 Redis node B
user:456 12431 Redis node C

This distributes the workload. You can inspect the slot for a key with:

redis-cli -c CLUSTER KEYSLOT "user:123"

The cluster maps keys to slots, and slots to nodes.

1.5 Hash tags

Sometimes you want related keys to live on the same shard:

user:{123}:profile
user:{123}:preferences
user:{123}:conversation

The {123} part is a hash tag. Redis hashes only the content inside the braces to choose the slot, so these related keys land on the same slot. That matters when you need multi-key operations:

keys = [
    "user:{123}:profile",
    "user:{123}:preferences",
]

Both can be colocated. But don't use hash tags everywhere: if every key uses {global}, you accidentally send everything to one slot and create a hotspot.

2. Hot keys

2.1 The hot-key problem

Imagine 10 million users, but one cache key, product:popular, is requested constantly:

  1. 10,000 requests/sec
  2. product:popular
  3. One shard

The cluster has 20 nodes, but one node becomes overloaded. This is called a hot key.

2.2 Hot keys in AI systems

AI applications can create extremely hot keys, for example system_prompt:v12, popular_embedding:abc, global_tool_schema:v8, popular_rag_query:xyz and model_config:v4.

Imagine 500,000 requests all needing the same system prompt, and so the same cache key. Even if the value is small, the request rate can overload a single cache node.

The problem isn't memory. It's request concentration.

2.3 Detecting hot keys

Track requests, bytes, hits and misses per key. For example:

hot_key_stats = {
    "prefix:company_policy:v12": 1_820_000,
    "embedding:abc": 12_500,
    "rag:query:xyz": 8_900,
}

Then identify the top 1%, 0.1% and 0.01% of keys by request volume. You may discover that one key is responsible for 35% of cache traffic. That is a very different scaling problem from uniformly distributed traffic.

3. Local + distributed caching

3.1 L1 local cache, L2 Redis

One of the best ways to reduce pressure on a distributed cache is to add a local cache in front of it:

  1. Application
  2. L1 local cacheprocess memory
  3. L2 Redison an L1 miss
  4. L3 database / model / RAGon an L2 miss

This creates a hierarchy. A simple Python L1 cache:

from cachetools import TTLCache
 
local_cache = TTLCache(maxsize=10_000, ttl=30)
 
 
def get_local(key):
    return local_cache.get(key)
 
 
def set_local(key, value):
    local_cache[key] = value

Now the application can avoid Redis entirely for very hot data.

3.2 L1 + L2 implementation

def get_cached(redis, key):
    # L1
    if key in local_cache:
        return local_cache[key]
 
    # L2
    value = redis.get(key)
    if value is not None:
        local_cache[key] = value
        return value
 
    return None
 
 
def set_cached(redis, key, value, ttl):
    local_cache[key] = value
    redis.set(key, value, ex=ttl)

3.3 But L1 creates a new problem

Suppose applications A, B and C each have their own local cache, and Redis holds version 2 of a key:

Copy Version
Redis 2
Application A 1
Application B 2
Application C 1

Now you have cache divergence. This is the trade-off:

L1 improves L1 complicates
Latency Invalidation
Redis load Consistency
Hot-key handling Memory usage

4. Invalidation at scale

4.1 Cache invalidation at scale

The classic problem: how do you invalidate cached data when the underlying data changes?

Suppose a document feeds an embedding cache, a retrieval cache and a response cache:

  1. Document v1
  2. Embedding cache
  3. Retrieval cache
  4. Response cache

When the document moves to v2, you potentially need to invalidate the document, embedding, retrieval, reranking, context and response caches. This is why dependency-aware cache design matters.

4.2 Version-based invalidation

Instead of deleting thousands of keys, change the version. knowledge:v1 becomes knowledge:v2, and the version is part of the key:

key = (
    f"rag:"
    f"tenant:{tenant_id}:"
    f"knowledge:{knowledge_version}:"
    f"query:{query_hash}"
)

Once knowledge_version = 2, every v1 entry becomes unreachable. You don't need to delete old keys synchronously. This is one of the safest invalidation strategies for AI systems.

4.3 Namespace versioning

A generalized approach:

def cache_key(namespace, version, identifier):
    return f"{namespace}:v{version}:{identifier}"
 
 
key = cache_key("rag", 12, query_hash)  # rag:v12:abc123

Deploy rag:v13, and rag:v12:abc123 is automatically ignored.

4.4 TTL is not invalidation

This distinction is extremely important:

Answers
TTL How long should this cache entry live?
Invalidation When is this cache entry no longer correct?

A 24-hour TTL doesn't mean the data stays correct for 24 hours. The underlying data might change after 10 seconds, or not for 30 days. Use versioning, events, explicit invalidation and TTLs for different purposes.

4.5 Event-driven invalidation

A production system can publish an event when data changes:

  1. Document updated
  2. Event bus
  3. Cache invalidation
  4. Embedding invalidated
  5. Retrieval invalidated
  6. Response invalidated

For example:

event = {
    "type": "document.updated",
    "tenant_id": "acme",
    "document_id": "doc_123",
    "version": 7,
}
 
 
def handle_document_update(event):
    document_id = event["document_id"]
    version = event["version"]
 
    invalidate_document(document_id)
    invalidate_embeddings(document_id)
    invalidate_retrieval(document_id, version)

At scale, events are usually better than trying to invalidate every dependent cache synchronously on the request path.

5. Cache stampedes

5.1 What is a cache stampede?

Consider a popular key with a 60-second TTL. At 12:00:00 the entry expires, and 50,000 requests arrive at once. All of them see a cache miss, and all of them call the backend:

  1. Popular key expires
  2. 50,000 cache misses
  3. 50,000 calls to the database or LLM

This is a cache stampede.

5.2 Why AI systems are especially vulnerable

Suppose the cache holds an expensive LLM response. One request costs one LLM call, say $0.05. Now 20,000 simultaneous misses can mean 20,000 LLM requests. You have created a cost explosion.

The cache was supposed to protect the model. Instead:

  1. Cache expiry
  2. Stampede
  3. LLM overload
  4. Latency
  5. Retries
  6. More load

5.3 Request coalescing

One solution: only one request should regenerate a missing value.

  1. 50 requests
  2. Cache miss
  3. Request leader
  4. Backend
  5. Cache
  6. Other 49 requests reuse the result

Conceptually:

async def get_or_generate(key):
    value = await cache.get(key)
    if value:
        return value
 
    lock = await acquire_lock(key)
 
    if lock:
        try:
            value = await cache.get(key)
            if value:
                return value
 
            value = await expensive_operation()
            await cache.set(key, value, ttl=300)
            return value
        finally:
            await release_lock(key)
 
    return await wait_for_value(key)

Notice the second cache lookup after acquiring the lock. That is important: another request may have filled the cache while we were waiting.

5.4 Distributed locks

If several application instances need to coordinate, a process-local lock isn't enough. You need distributed coordination.

Redis provides patterns for distributed locks, including the Redlock algorithm, but distributed locking has important correctness caveats and should be chosen carefully for the workload. A simple Redis lock pattern:

import uuid
 
lock_id = str(uuid.uuid4())
 
acquired = redis.set("lock:cache:key", lock_id, nx=True, ex=10)

NX means "only create the key if it doesn't exist", so only one worker becomes the leader.

5.5 Never delete someone else's lock

Suppose worker A gets the lock, and the lock expires before A is done. Worker B then gets the same lock. When A finally finishes, if it blindly runs:

redis.delete("lock:cache:key")

it deletes worker B's lock.

So the lock's value should identify its owner (a random token), and a worker should release the lock only if the token still belongs to it. This is why production lock implementations use an atomic check-and-delete rather than an unconditional delete. Redis's own documentation discusses this pattern, and the care needed around lock validity and fencing.

5.6 Stale-while-revalidate

Another powerful strategy: don't treat expired data as unusable straight away.

  1. Request
  2. Cache entry
  3. Freshreturn itStale but acceptablereturn it, refresh in background

For example:

async def get_swr(key):
    entry = await cache.get(key)
 
    if entry is None:
        return await regenerate(key)
 
    if entry.is_fresh:
        return entry.value
 
    if entry.is_within_stale_window:
        trigger_background_refresh(key)
        return entry.value
 
    return await regenerate(key)

This is extremely useful for RAG responses, frequently requested embeddings, product metadata, public knowledge, model configuration and expensive API responses.

6. Warming and negative caching

6.1 Cache warming

Sometimes you know what users will request. Every morning, your application may get traffic for popular reports, documents, embeddings and RAG questions. Instead of making the first user pay for the cache miss, precompute:

  1. User
  2. Cache miss
  3. Expensive computation
Without warming
  1. Scheduler
  2. Generate cache entries
  3. Store
  4. Users receive cached results
With warming

For example:

async def warm_cache(keys):
    for key in keys:
        if await redis.exists(key):
            continue
 
        value = await expensive_operation(key)
        await redis.set(key, value, ex=3600)

6.2 Cache warming for AI systems

Useful candidates include popular embeddings, popular retrieval queries, popular product answers, common system prompts, frequently used tool metadata, popular model configuration and common agent workflows.

But avoid warming everything. Otherwise you spend GPU, CPU, API cost and storage generating data nobody uses. Warm only entries that are high probability + high cost + high impact.

6.3 Negative caching

Caching successful results is obvious. But sometimes you should cache failures or "not found" results too.

A user asks for document XYZ and the system finds nothing. Without negative caching, every repeat of that request goes to the database. Instead, cache the absence with a short TTL:

await redis.set("document:XYZ", "__NOT_FOUND__", ex=30)

Now repeated invalid requests don't hammer your database. In AI systems, this also applies to no retrieval results, no matching tool, no model route and no user configuration.

Use short TTLs, because absence can change.

7. Tenants and access control

7.1 Multi-tenant caching

AI SaaS applications serve many tenants. Never create a key like rag:query:<hash> if the result depends on tenant-specific data. Instead:

key = f"tenant:{tenant_id}:rag:{query_hash}"

Otherwise:

  1. Tenant A's result
  2. Cache
  3. Served to Tenant B

That's not a performance bug. That's a security incident.

7.2 ACLs must still be applied

Even with tenant_id in the key, don't treat the cache as your authorization layer. The flow should remain:

  1. Authenticate
  2. Authorize
  3. Build cache scope
  4. Look up cache
  5. Return permitted data

For RAG:

  1. User
  2. ACL filters
  3. Retrieve permitted documents
  4. Cache the result, scoped appropriately

A cache hit should never bypass your access-control model.

8. Multi-region and consistency

8.1 Multi-region caching

Now imagine users in India, the US, Europe and Singapore. A single Redis cluster means high network latency for most of them. Instead, each region gets its own cache:

  1. Global users
  2. IndiaUSEurope
  3. RedisRedisRedis
  4. Data systems

The closest region serves the request. This reduces network latency, cross-region traffic and dependence on a single region.

8.2 Locality matters

Suppose a user in Bengaluru calls a Redis cluster in the US. Even if Redis responds quickly, the network round trip is still there.

For high-throughput AI applications, 5 ms may matter. For a 2-second LLM generation, 5 ms may not matter much. So regional caching should be driven by your latency budget, request volume, data locality, consistency requirements and cost, not simply because "multi-region sounds scalable".

8.3 Global vs regional data

Not all cache data should be replicated globally:

Scope Candidates
Global Model configuration, public documentation, global feature flags, static system prompts
Regional Session data, user-specific context, regional API responses, localized RAG results
User-local Conversation state, personalization, private memory

Think in widening circles:

  1. Global
  2. Regional
  3. Tenant
  4. User
  5. Session

The narrower the scope, the less likely the data should be replicated globally.

8.4 Read replicas

For read-heavy workloads, replicas can spread the read load:

  1. Primary
  2. Replica 1Replica 2Replica 3

Redis replication keeps replicas following the primary's dataset. But a replica is not automatically strongly consistent: there can be replication lag. Don't use replica reads for correctness-sensitive decisions without understanding the consistency guarantees of your deployment.

8.5 Cache consistency models

Different applications need different consistency:

Model Meaning Examples
Strong consistency Must be correct immediately Authorization, financial state, critical configuration
Eventual consistency An old value can be served briefly Popular product metadata, public recommendations, non-critical analytics
Stale-acceptable Slightly stale data is allowed on purpose, for lower latency Popular content, aggregated statistics, non-critical AI recommendations

The cache strategy should follow the business requirement, not the other way around.

9. Memory, eviction and TTLs

9.1 Cache eviction

Eventually, cache memory fills up and you need an eviction policy. Common approaches include LRU, LFU, TTL-based, random and no eviction. Redis supports several eviction policies, including LRU- and LFU-style strategies, for managing memory pressure.

9.2 LRU vs LFU

Policy Removes Good when
LRU (least recently used) Entries not accessed recently Recent activity predicts future activity
LFU (least frequently used) Entries accessed least often Some keys are consistently popular

For AI workloads, LFU can be interesting when you have long-lived hot prompts, popular RAG queries or popular product questions. But always benchmark against your actual access pattern.

9.3 TTL jitter

Imagine 1 million cache entries created at 12:00:00, all with a 1-hour TTL. At 13:00:00 they all expire together, creating a massive spike.

Add jitter instead:

import random
 
ttl = 3600 + random.randint(-300, 300)

Now expirations are spread across a 10-minute window instead of landing in a single second. It's a tiny implementation detail with significant production value.

9.4 TTL should reflect data volatility

Don't use a 1-hour TTL for everything. Ask how often the underlying data changes:

Data TTL
Product availability 30 seconds
Popular public content 10 minutes
Session state 15 minutes
Model metadata 24 hours
Embeddings Potentially long-lived, and versioned
Authorization data Much stricter

There is no universal TTL.

10. Payloads

10.1 Cache payload size matters

A cache isn't free. 10 million objects at 100 KB each is roughly 1 TB, before overhead and replication.

In AI systems, payloads can become huge: large RAG contexts, tool outputs, documents, conversation histories, serialized embeddings and model metadata. Ask:

Should I cache the entire object, or only the expensive part?

10.2 Cache IDs instead of large objects

Instead of caching the full document contents under a retrieval key, cache the chunk IDs and fetch the content from the document store. Similarly, an agent result might cache an artifact_id instead of a massive artifact.

This reduces cache memory.

10.3 Compression

For large payloads, compress before storing:

import gzip
import json
 
 
def encode(value):
    raw = json.dumps(value).encode()
    return gzip.compress(raw)
 
 
def decode(value):
    raw = gzip.decompress(value)
    return json.loads(raw)

But compression isn't automatically beneficial. It adds CPU cost, so measure what you save in network transfer and memory against the CPU spent compressing and decompressing.

10.4 Serialization matters

JSON is convenient, but it can be larger and slower than binary formats. For high-throughput systems, consider MessagePack, Protobuf, Arrow or a custom binary representation, depending on the payload.

The right question: what is the smallest representation that is fast enough to serialize and deserialize?

11. Deciding what to cache

11.1 Don't cache everything

One of the biggest mistakes is "if it's expensive, cache it". Not necessarily. A cache entry should have high reuse + high computation cost + acceptable staleness + a reasonable memory footprint.

A useful mental model:

Cache value ≈ reuse frequency × cost avoided × correctness value

If something is only ever used once, caching it may not help.

11.2 Cache admission

Instead of caching every expensive result, cache results that are likely to repeat:

def should_cache(request_count, compute_cost, payload_size):
    return (
        request_count >= 3
        and compute_cost >= 0.5
        and payload_size < 1_000_000
    )

A production implementation would use historical access patterns rather than simple thresholds.

11.3 Adaptive TTL

You can even make the TTL depend on popularity:

def choose_ttl(hit_count):
    if hit_count > 10_000:
        return 3600
 
    if hit_count > 1_000:
        return 900
 
    return 120

Popular objects get longer TTLs; rare objects get shorter ones. This can improve cache efficiency.

11.4 Cache warming + adaptive TTL

Now combine them:

  1. High hit count
  2. Long TTL
  3. Pre-warm
  4. L1 + L2 cache
Popular key
  1. Low hit count
  2. Short TTL
  3. Don't pre-warm
Rare key

This is much more efficient than applying one caching policy to every request.

12. Failure handling

12.1 A production cache will fail

Plan for Redis timeouts, Redis being unavailable, connection-pool exhaustion, cluster failover, network partitions, memory exhaustion, serialization errors, corrupt values, stale values and hot shards.

Your application has to decide: what happens if the cache is unavailable?

12.2 Cache-aside

The most common pattern:

  1. Application
  2. Cache
  3. Databaseon a miss
  4. Cachestore the result

In code:

def get_user(user_id):
    key = f"user:{user_id}"
 
    value = redis.get(key)
    if value:
        return deserialize(value)
 
    value = database.get_user(user_id)
    redis.set(key, serialize(value), ex=300)
    return value

And if Redis fails, fall back to the source:

def get_user(user_id):
    key = f"user:{user_id}"
 
    try:
        value = redis.get(key)
        if value:
            return deserialize(value)
    except Exception:
        pass
 
    return database.get_user(user_id)

This is graceful degradation.

12.3 Fail open vs fail closed

For caches, this distinction matters:

Mode If the cache fails Suits
Fail open Continue to the source Non-critical acceleration
Fail closed Reject the operation Potentially security-sensitive state

Don't choose based on technical convenience. Choose based on business correctness, security and data sensitivity.

12.4 Timeouts matter

Never let redis.get(key) block your application indefinitely. Use bounded timeouts:

redis_client = Redis(
    host="cache",
    socket_connect_timeout=0.05,
    socket_timeout=0.05,
)

The exact values depend on your latency budget. The principle:

A cache must not become a new source of indefinite latency.

12.5 Connection pooling

High-throughput applications shouldn't open a new Redis connection for every request. Use a connection pool:

from redis import Redis
from redis.connection import ConnectionPool
 
pool = ConnectionPool(host="localhost", port=6379, max_connections=100)
 
redis = Redis(connection_pool=pool)

This reduces connection setup, TCP overhead and resource churn. But don't blindly set max_connections = 10_000, or your application can overwhelm Redis.

12.6 Backpressure

Suppose your application sends 100,000 requests per second, but Redis can safely handle 50,000. You need backpressure. Otherwise:

  1. Request rate
  2. Redis overload
  3. Latency
  4. Timeouts
  5. Retries
  6. Even more traffic

Retries without limits can become a self-inflicted denial of service.

12.7 Retry carefully

Bad:

for _ in range(100):
    retry()

Better:

from time import sleep
 
from redis.exceptions import TimeoutError
 
for attempt in range(3):
    try:
        return redis.get(key)
    except TimeoutError:
        if attempt == 2:
            break
        sleep(0.01 * (2**attempt))

Combine exponential backoff + jitter + a maximum number of retries + a timeout.

12.8 Retry storms

Imagine Redis goes down and every request retries. Your traffic is now the original traffic plus all the retry traffic, which makes recovery harder.

A production system should use bounded retries + circuit breakers + backoff + a fallback.

12.9 Circuit breaker

Conceptually:

  1. Redis health check
  2. Healthynormal operationFailingopen the circuit, stop calling Redis for a while, use the fallback

For example:

class CircuitBreaker:
    def __init__(self, threshold=5):
        self.failures = 0
        self.threshold = threshold
        self.open = False
 
    def failure(self):
        self.failures += 1
        if self.failures >= self.threshold:
            self.open = True
 
    def success(self):
        self.failures = 0
        self.open = False

A real implementation also needs a cooldown, a half-open state, time windows and concurrency control.

13. Observability and cost

13.1 Observability

You can't operate distributed caching without metrics. At minimum:

Area Metrics
Effectiveness Hit rate, miss rate
Latency Cache latency, P95, P99
Throughput Requests/sec, bytes/sec
Memory Usage, fragmentation, evictions, expired keys
Hotspots Hot keys, hot shards
Connections Pool usage, timeouts, errors
Cluster Replication lag, cluster health, failovers
Coordination Stampede events, lock contention

13.2 AI-specific metrics

For AI workloads, add LLM tokens avoided, embedding calls avoided, RAG searches avoided, reranker calls avoided, tool calls avoided, GPU compute avoided and estimated cost avoided:

metrics.increment("ai.cache.tokens_saved", value=12_450)
metrics.increment("ai.cache.embedding_calls_saved")
metrics.increment("ai.cache.tool_calls_saved")

Now the cache isn't just a technical component. You can quantify its business value.

13.3 Cache cost accounting

Imagine:

Item Monthly
LLM spend $100,000
Saved by caching $28,000
Redis cost $4,000
Net saving $24,000

That is the metric leadership actually cares about. Not "Redis hit rate = 83%", but cache infrastructure cost vs compute cost avoided.

14. Dependencies and versioned deployments

14.1 The cache dependency graph

By now, your AI application may have a chain like this, with a cache at every layer:

  1. Document
  2. Embedding
  3. Retrieval
  4. Reranking
  5. Context
  6. Prompt
  7. Prefix / KV
  8. Response

If document v1 changes, potentially everything downstream of it is invalid. This is why cache invalidation gets harder as your AI architecture gets richer.

14.2 Dependency-aware versioning

Instead of deleting everything by hand, make the document, embedding, index, reranker, prompt and model versions part of the cache identity:

key = (
    f"response:"
    f"tenant:{tenant_id}:"
    f"model:{model_version}:"
    f"prompt:{prompt_version}:"
    f"index:{index_version}:"
    f"query:{query_hash}"
)

Then a change to index_version automatically creates a new cache namespace.

14.3 The cache dependency matrix

A useful production design exercise:

Cache Depends on Invalidate when
Embedding Text, model Model or preprocessing changes
Retrieval Query, embedding, index Index changes
Reranking Query, candidates, reranker Candidates or model change
Context Retrieval, policy Source or policy changes
Response Context, model, prompt Any dependency changes
Prefix Model, tokenizer, template, prompt Any prefix dependency changes
Tool result Tool inputs, tool version Tool or data changes

This makes invalidation explicit.

14.4 Cache namespaces

A large system should use namespaces, for example ai:embedding, ai:retrieval, ai:rerank, ai:response, ai:agent, ai:tool and ai:prefix, each with its own version:

ai:embedding:v3:<hash>
ai:retrieval:v8:<hash>
ai:response:v12:<hash>
ai:prefix:v5:<hash>

This makes debugging, monitoring, migration and invalidation much easier.

14.5 Blue/green cache migration

Suppose you're moving from model v4 to model v5. Don't overwrite model:v4 immediately. Run model:v5 alongside it and shift traffic gradually:

Stage v4 v5
1 90% 10%
2 50% 50%
3 0% 100%

Old cache data expires naturally. This dramatically reduces migration risk.

14.6 Cache stampede + version deployment

A subtle problem: you deploy prompt:v13, and suddenly every request misses, because every key changed from prompt:v12. You have created an artificial cold start.

Mitigations include cache warming, gradual rollout, dual reads and background population. For example:

  1. v12 active, v13 warming
  2. v13 becomes active

14.7 Dual read during migration

During a migration:

value = cache.get(new_key)
 
if value is None:
    value = cache.get(old_key)
 
    if value is not None:
        cache.set(new_key, value, ex=ttl)

This can avoid an immediately cold cache. But only do it when the old value is actually compatible with the new system. Never use migration convenience as an excuse to break cache correctness.

15. The cache hierarchy

15.1 A multi-level AI cache

A mature architecture might look like:

  1. User
  2. Application L1
  3. Redis L2on an L1 miss
  4. RAG cacheon a miss
  5. Embedding / retrievalon a miss
  6. Prompt prefixon a miss
  7. LLM / GPU

The higher levels are cheaper, faster and smaller. The lower levels are more expensive, slower and more computationally valuable.

15.2 Every hit prevents work further down

Level What it is
L1 Process memory
L2 Redis
L3 Database / vector DB
L4 External APIs
L5 LLM / GPU

Every cache hit prevents work further down the stack. So an L1 hit is often better than an L2 hit, which is better than an L3 hit, which is better than an LLM call.

15.3 But more cache layers mean more complexity

This is the trap. You might start with just Redis, then add a local cache, a RAG cache, an embedding cache, a response cache and a prefix cache. Now you have six cache layers, and each can go stale independently.

Add a cache only when its benefit justifies the operational complexity.

16. A production cache abstraction

16.1 One interface

You can standardize application-level caching behind one interface:

class Cache:
    async def get(self, key):
        raise NotImplementedError
 
    async def set(self, key, value, ttl):
        raise NotImplementedError
 
    async def delete(self, key):
        raise NotImplementedError
 
 
class RedisCache(Cache):
    def __init__(self, redis):
        self.redis = redis
 
    async def get(self, key):
        return await self.redis.get(key)
 
    async def set(self, key, value, ttl):
        await self.redis.set(key, value, ex=ttl)
 
    async def delete(self, key):
        await self.redis.delete(key)

Now your application doesn't need to know whether the backend is Redis, Memcached, local memory, a cloud cache or a mock cache.

16.2 Add observability to the abstraction

import time
 
 
class ObservableCache:
    def __init__(self, backend, metrics):
        self.backend = backend
        self.metrics = metrics
 
    async def get(self, key):
        start = time.monotonic()
 
        try:
            value = await self.backend.get(key)
 
            if value is None:
                self.metrics.increment("cache.miss")
            else:
                self.metrics.increment("cache.hit")
 
            return value
        finally:
            self.metrics.observe("cache.latency", time.monotonic() - start)

Now every cache operation automatically produces metrics.

16.3 Cache policy as configuration

Don't hardcode caching behaviour everywhere. Instead:

CACHE_POLICIES = {
    "embedding": {"ttl": 86400, "compression": True},
    "rag_result": {"ttl": 300, "compression": True},
    "response": {"ttl": 60, "compression": True},
    "session": {"ttl": 1800, "compression": False},
}
 
policy = CACHE_POLICIES["embedding"]
 
await cache.set(key, value, ttl=policy["ttl"])

This makes caching behaviour centrally configurable.

17. Putting it together

17.1 Production cache architecture

  1. Users
  2. API gateway
  3. Auth / rate limit
  4. Application
  5. L1 local cacheRequest router
  6. Redis Cluster
  7. Shard A+ replicaShard B+ replicaShard C+ replica
  8. Cache miss
  9. RAGToolsDatabase
  10. Prompt / prefix
  11. LLM engine
  12. GPU

This is no longer "just Redis". It is a distributed caching architecture.

17.2 Production failure scenarios

You should test these explicitly:

Scenario The question
Redis completely unavailable What happens?
One Redis shard unavailable What happens?
Replica lag Can stale data be served?
Cache becomes full Which entries are evicted?
A hot key appears Can one shard handle it?
100K simultaneous cache misses Does the backend survive?
Network latency increases Does application latency explode?
New model deployment Does the old cache remain compatible?
Tenant data changes Are dependent caches invalidated?
A region goes down Can traffic fail over?

17.3 Chaos testing the cache

Treat a production cache as a failure-prone dependency. Kill Redis, delay Redis, drop connections, force cache misses, fill cache memory, create hot keys, expire popular keys at the same time, introduce replication lag, restart nodes and move shards. Then measure: does the application survive?

The goal isn't that Redis never fails. The goal is that Redis can fail without taking down the AI system.

18. The full picture

18.1 The golden rules of distributed AI caching

  1. A cache is disposable. When correctness matters, the source of truth lives somewhere else.
  2. Every cache key needs the right scope: tenant, user, model, version, permissions.
  3. Version dependencies instead of relying only on TTLs.
  4. Never let cache misses overwhelm your backend. Use request coalescing, locks, stale-while-revalidate and warming.
  5. Monitor hot keys.
  6. Design for cache failure.
  7. Don't confuse replication with strong consistency.
  8. Don't add cache layers without measuring their value.
  9. Track compute avoided, not just cache hits.
  10. Treat cached data as potentially sensitive.

18.2 The final production mental model

When designing caching for an AI system, think in this order:

  1. What expensive work can be reused?
  2. What is the cache scope?
  3. What makes the result invalid?
  4. What should the key contain?
  5. How long should it live?
  6. What happens on a miss?
  7. What happens during a stampede?
  8. What happens when Redis fails?
  9. How does the cache scale?
  10. How do we observe its value?

This mindset matters much more than knowing any particular Redis command.

18.3 The complete AI Caching Playbook

We've gone from "What is caching?" to "How do I architect caching across an entire AI platform?". The layers now look like:

  1. AI platform
  2. Applicationresponse, session, memoryRAGembedding, retrieval, rerankAgentstool, state, workflow
  3. Prompt / prefix
  4. KV cache
  5. GPU cache
  6. Distributed cache
  7. L1localL2RedisL3database

The final architecture is not about one cache. It's about controlled reuse across the entire computation graph.

18.4 Part 7 takeaways

If you remember only these:

  1. Distributed caching is a systems problem. Once you scale, you need sharding, replication, failover, routing and observability.
  2. Redis Cluster distributes keys using hash slots: 16,384 of them, spread across the nodes.
  3. Hot keys can break an otherwise well-scaled cluster. One extremely popular key can overload one shard.
  4. L1 + L2 caching can dramatically reduce cache pressure: local memory, then Redis, then the source.
  5. TTL isn't invalidation. Combine TTLs with versioning and events where appropriate.
  6. Cache stampedes are dangerous. Use request coalescing, distributed coordination, stale-while-revalidate, cache warming and TTL jitter.
  7. Cache keys define correctness. A bad key can create stale data, cross-tenant leaks, wrong-model results and incorrect RAG answers.
  8. Cache failure must be survivable. A cache should speed up your system, not become a single point of failure.
  9. Distributed locks need careful correctness design. SET NX alone doesn't solve every coordination problem; lock expiry, ownership, fencing and failure modes matter.
  10. Measure business value. Track tokens avoided, GPU compute avoided, API calls avoided, latency saved and cost saved, not just the cache hit rate.

18.5 The core principle

After seven parts, the central idea fits in one sentence:

A production AI cache should make expensive work reusable without compromising correctness, security, scalability or reliability.

That's the difference between "I added Redis" and "I designed a production caching architecture". The first is infrastructure. The second is engineering.

What's next?

Part 8: Cache invalidation, consistency and correctness in AI systems

We've covered how to build and scale caches. But there is one problem developers have feared for decades:

What happens when the cache is wrong?

Part 8 goes deeper into:

  • cache invalidation strategies, TTL vs versioning, and generational caching
  • write-through, write-behind, cache-aside and read-through caching
  • event-driven and dependency-based invalidation, and versioned cache keys
  • strong vs eventual consistency, stale data and cache coherence
  • race conditions, concurrent writes and lost updates
  • distributed locks, fencing tokens and optimistic concurrency
  • RAG knowledge updates, embedding invalidation and agent memory invalidation
  • conversation state consistency and model/prompt version invalidation
  • multi-tenant correctness and security implications
  • testing stale-cache scenarios and production invalidation architecture

Because the hardest caching question isn't "How do I make it fast?". It's:

How do I make it fast without returning the wrong answer?

And in AI systems, wrong cached information can be far more dangerous than a slow request.

Next: Part 8, Cache invalidation and correctness covers TTLs, versioned keys, dependency graphs, events and testing for correct caches.