HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 9: Cache Observability and Cost

Measure what the cache actually saves: latency, tokens, calls avoided, cost, ROI, hit quality and freshness, not just the hit rate.

Haribaskar Dhanabalan16 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

Measure what the cache actually saves.

This is Part 9 of the AI Caching Playbook. The earlier parts:

A cache can have a 95% hit rate and still be a terrible cache. Hit rate alone doesn't tell you:

  • whether the cached result was correct
  • whether it was fresh
  • how much latency was actually saved
  • how many tokens were avoided
  • how much money was saved
  • whether the hits are concentrated on a tiny number of requests
  • whether the cache costs more infrastructure than it saves

In AI systems, caching isn't just an infrastructure optimization. It is an economic optimization. You are trying to reduce latency, LLM calls, embedding calls, vector searches, reranker calls, tool calls, GPU computation and infrastructure cost, while keeping answer quality, freshness, correctness and security.

So the real question isn't "How many requests hit the cache?". It is:

How much useful work did the cache prevent?

1. Why cache hit rate is not enough

1.1 Hit rate vs value

The most common cache metric is:

Cache hit rate = cache hits ÷ total requests

With 1,000 requests and 900 hits, the hit rate is 90%. That looks excellent. But compare two caches:

Hits Saved per hit Total saved
Cache A 900 $0.0001 $0.09
Cache B 500 $0.08 $40

If Cache A costs $50 a month to run, it may not be economically useful at all. Cache B, with far fewer hits, produces much more value.

So production AI systems should measure hit rate + latency saved + tokens saved + compute saved + calls avoided + cost saved + quality preserved.

2. Capturing cache events

2.1 Start with a cache event

Before building dashboards, capture cache events:

from dataclasses import dataclass, field
from time import time
 
 
@dataclass
class CacheEvent:
    cache_name: str
    key: str
    hit: bool
    lookup_ms: float
    saved_tokens: int = 0
    saved_cost: float = 0.0
    source: str = "unknown"
    timestamp: float = field(default_factory=time)

A cache lookup then produces an event:

event = CacheEvent(
    cache_name="embedding_cache",
    key=cache_key,
    hit=True,
    lookup_ms=2.4,
    saved_tokens=0,
    saved_cost=0.0002,
    source="redis",
)

This event becomes the foundation for observability.

2.2 Measure hits and misses

The first layer is basic:

cache_hits = 0
cache_misses = 0
 
 
def record_cache_result(hit: bool):
    global cache_hits, cache_misses
 
    if hit:
        cache_hits += 1
    else:
        cache_misses += 1
 
 
def hit_rate():
    total = cache_hits + cache_misses
    if total == 0:
        return 0
    return cache_hits / total

But don't stop there. Track it per cache type: embedding, retrieval, reranker, response, tool, agent state and prefix caches. A single global hit rate hides too much.

2.3 Measure hit rate by route

Suppose your application has /chat, /rag, /search, /agent and /summarize. The global hit rate might be 82%, but per route:

Route Hit rate
/chat 95%
/rag 88%
/search 72%
/agent 21%

That tells a very different story. Record dimensions alongside each result:

metric = {
    "cache": "response_cache",
    "route": "/rag",
    "model": "model-x",
    "tenant": "tenant_123",
    "hit": True,
}

Now you can see where caching is actually effective.

2.4 Don't put high-cardinality data into every metric

This is an important production concern. Don't create metrics like cache_hit{user_id="123456"} for millions of users: that creates enormous metric cardinality.

Use low-cardinality labels instead, such as cache_hit{cache="response", route="/rag", model="model-x"}, and keep the detailed information in logs and traces rather than metric labels.

3. Latency

3.1 Measure cache lookup latency

Caching adds its own latency. A 2 ms Redis lookup in front of an 80 ms vector query and a 1,200 ms LLM generation is clearly useful. But if the cache lookup itself takes 250 ms, you have a problem.

import time
 
 
def lookup_cache(cache, key):
    start = time.perf_counter()
    value = cache.get(key)
    latency_ms = (time.perf_counter() - start) * 1000
    return value, latency_ms

Track the lookup latency at P50, P95 and P99.

3.2 Measure end-to-end latency saved

The most useful latency metric isn't "cache lookup = 2 ms". It's the original request latency minus the cached request latency:

Latency
Without cache 1,850 ms
With cache 35 ms
Saved 1,815 ms
latency_saved_ms = uncached_latency_ms - cached_latency_ms

Then aggregate the average, P50 and P95 latency saved.

4. Work avoided

4.1 Token savings

For LLM applications, tokens are one of the most important metrics. If a request normally needs 5,000 input tokens and 1,000 output tokens, a response cache hit avoids 6,000 tokens:

saved_tokens = uncached_input_tokens + uncached_output_tokens

Aggregate the total tokens saved, tokens saved per request and tokens saved per cache.

4.2 Measure savings by cache type

Different caches prevent different work:

Cache What it saves
Response cache Input tokens + output tokens
Prompt / prefix cache Repeated prefill computation
Embedding cache Embedding-model calls
Retrieval cache Vector database work
Tool cache External API calls

So don't combine everything into one tokens_saved number. Track response_cache_tokens_saved, embedding_calls_avoided, retrieval_calls_avoided and tool_calls_avoided separately.

4.3 LLM calls avoided

This is more useful than cache hits. With 10,000 requests and 6,000 response cache hits, you avoided 6,000 LLM calls:

llm_calls_avoided += 1

Now you can answer "How many expensive model executions did our cache eliminate?", which means much more than "Our cache has an 85% hit rate".

4.4 Embedding calls avoided

In a RAG system, a query is normally embedded before the vector search. If the query embedding is cached, you skip the embedding-model call:

  1. User query
  2. Embedding cache hit
  3. Vector search

Track both the calls and the tokens:

embedding_calls_avoided += 1
embedding_tokens_avoided += token_count

This lets you calculate the actual savings.

4.5 Tool calls avoided

Agentic systems can have expensive tools: CRM, payment and search APIs, databases, weather APIs and internal services. If an agent makes 100,000 tool calls a day and caching avoids 30,000 of them, that's significant.

tool_call_event = {
    "tool": "customer_search",
    "cache_hit": True,
    "call_avoided": True,
}

From these events you can calculate the calls, latency and external API cost avoided.

5. Cost and ROI

5.1 Cost saved

Now the important part. If an LLM request costs $0.02 and caching avoided 5,000 requests:

llm_cost_saved = 5000 * 0.02  # $100

Do the same for embeddings, reranking, vector DB queries, external APIs, GPU compute and database queries.

5.2 Build a cost model

Don't hardcode cost calculations throughout your application. Centralize them:

COSTS = {
    "llm_request": 0.02,
    "embedding_request": 0.0003,
    "reranker_request": 0.001,
    "tool_call": 0.005,
}
 
 
def calculate_savings(llm_calls=0, embedding_calls=0, reranker_calls=0, tool_calls=0):
    return (
        llm_calls * COSTS["llm_request"]
        + embedding_calls * COSTS["embedding_request"]
        + reranker_calls * COSTS["reranker_request"]
        + tool_calls * COSTS["tool_call"]
    )

5.3 Real cost models should use tokens

LLM costs usually depend on tokens. A more realistic model:

def llm_cost(input_tokens, output_tokens, input_price_per_million, output_price_per_million):
    return (
        input_tokens / 1_000_000 * input_price_per_million
        + output_tokens / 1_000_000 * output_price_per_million
    )
 
 
cost = llm_cost(
    input_tokens=5000,
    output_tokens=1000,
    input_price_per_million=5,
    output_price_per_million=15,
)  # $0.04

This lets your observability system estimate the value of each cache hit.

5.4 Cache ROI

Caching has infrastructure costs too. If the cache costs $300 a month and saves $2,400 a month, the net saving is $2,100. A simple ROI calculation:

def cache_roi(savings, cache_cost):
    if cache_cost == 0:
        return float("inf")
 
    return (savings - cache_cost) / cache_cost

In this example, ROI = 7x. That is a much better business metric than the hit rate.

6. Memory and infrastructure health

6.1 Cache memory efficiency

A cache can become expensive because of memory. Track used, allocated and free memory, evictions, entry count and average entry size:

memory_efficiency = used_memory_bytes / allocated_memory_bytes

But memory utilization alone isn't enough. You also want savings per GB:

Memory Savings / month
Cache A 100 GB $500
Cache B 20 GB $800

Cache B is dramatically more efficient.

6.2 Eviction rate

If entries keep getting evicted before they're reused, you're wasting memory. Track evictions per second, hour and day, and the eviction rate:

eviction_rate = evictions / total_cache_operations

A high eviction rate may mean the cache is too small, TTLs are too long, the cache policy is poor, the cache is polluted, objects are huge, or your workload assumptions are wrong.

6.3 Hot-key detection

Some entries receive enormous traffic: key A gets 100,000 requests while key B gets 50 and key C gets 2. Key A is a hot key, and it can cause Redis CPU pressure, network bottlenecks, single-node pressure and lock contention.

from collections import Counter
 
key_usage = Counter()
 
 
def record_access(key):
    key_usage[key] += 1

In production, use sampled or aggregated tracking rather than keeping unbounded per-key data in application memory.

6.4 Cache stampede detection

Imagine:

  1. Cache entry expires
  2. 10,000 requests arrive
  3. 10,000 cache misses
  4. 10,000 LLM calls

Your cache just created a disaster. Track misses per key within a time window:

if misses_for_key > THRESHOLD:
    alert("Possible cache stampede")

Production systems combine this with request coalescing, single-flight locks, stale-while-revalidate and probabilistic early refresh.

6.5 Cache penetration

Cache penetration happens when requests keep asking for data that isn't cacheable or doesn't exist. A user asks for document X: cache miss, database miss. The same request comes again: cache miss, database miss. You keep hitting the expensive backend.

One solution is a negative cache:

NOT_FOUND = "__NOT_FOUND__"
 
cache.set(key, NOT_FOUND, ttl=30)

But be careful: negative caching also needs correct invalidation.

6.6 Cache pollution

Not every result deserves to be cached. With 10,000 unique, one-time queries, caching all of them fills your cache with data that will never be reused. This is cache pollution. A simple admission rule:

def should_cache(frequency, generation_cost, result_size):
    return (
        frequency >= 2
        and generation_cost > 0.01
        and result_size < 1_000_000
    )

The exact thresholds depend on your workload. The principle:

Cache valuable work, not everything.

7. Quality

7.1 Semantic cache quality

Semantic caches introduce a new problem. "What is our refund policy?" and "Can customers get their money back?" may be semantically similar, but similarity doesn't guarantee identical answers. A semantic cache hit therefore needs evaluation:

def evaluate_cached_answer(query, cached_answer, current_context):
    # Application-specific evaluation.
    return {
        "relevant": True,
        "grounded": True,
        "fresh": True,
    }

Sample semantic cache hits and evaluate their relevance, groundedness, freshness and correctness.

7.2 RAG cache evaluation

RAG adds another dimension: did the cache return the right evidence? A retrieval cache might return chunk_10, chunk_20 and chunk_50 after the knowledge base has changed. The hit rate stays high while the evidence is obsolete.

So track the retrieval cache hit together with an index version match, a document version match and answer quality. A useful evaluation pipeline:

  1. Cache hit
  2. Sample
  3. Freshness check
  4. Groundedness check
  5. Quality evaluation

7.3 Shadow cache

A powerful production technique is the shadow cache. The application behaves as if there were no cache, but you look up the cache at the same time, without using its answer:

  1. Request
  2. Normal generationwhat users getShadow cache lookuprecorded, never served
  3. Compare the two

Then compare what the cache would have returned with what the system actually generated:

cached = cache.get(key)
 
fresh = generate_response(request)
 
if cached:
    compare(cached_answer=cached, fresh_answer=fresh)

This is extremely useful before enabling a semantic or response cache. You can estimate the potential hit rate, potential cost savings and potential quality loss without risking production correctness.

7.4 Sample cache hits for evaluation

Don't evaluate every cache hit with another LLM; that defeats the purpose. Sample instead:

import random
 
 
def should_sample(rate=0.01):
    return random.random() < rate

For example, evaluate 1% of cache hits, and track sampled, correct, incorrect and stale hits. Then estimate the cache's quality from the sample.

7.5 Cache hit quality

Define:

Hit quality = correct cache hits ÷ evaluated cache hits

If 970 of 1,000 sampled hits are correct, hit quality is 97%. Now you have two independent metrics, for example a hit rate of 85% and a hit quality of 97%. That is far more useful than either alone.

7.6 A cache effectiveness score

You can build an internal score that combines several dimensions:

def cache_effectiveness(hit_rate, quality, latency_saved_ratio, cost_saved_ratio):
    return hit_rate * quality * latency_saved_ratio * cost_saved_ratio

This isn't a universal industry metric. It's an internal decision-making metric. The important idea:

Evaluate a cache on useful work avoided, not on cache activity alone.

8. Proving the value

8.1 Compare before and after

Always measure the system without the cache first:

Before After
P95 latency 1,800 ms 420 ms
LLM calls 100K 61K
Embedding calls 100K 72K
Tokens 500M 310M
Cost $8,000 $4,700
Quality 91% 91%

This is the real cache story. Not "cache hit rate: 82%".

8.2 A/B test caching

For high-risk response caches, roll out gradually: for example, 90% of users with no cache and 10% with the cache enabled. Compare latency, cost, quality, errors, freshness and user satisfaction.

Assign users to buckets with a stable hash:

import hashlib
 
 
def cache_enabled(user_id):
    bucket = int(hashlib.sha256(str(user_id).encode()).hexdigest(), 16) % 100
    return bucket < 10

This gives roughly a 10% treatment group and a 90% control group. Use a stable hash such as SHA-256 rather than Python's built-in hash(), which is randomized per process for strings, so the same user could land in different groups on different servers.

9. Instrumentation

9.1 A cache observability event schema

A practical event might look like:

event = {
    "timestamp": "2026-10-08T12:00:00Z",
    "cache": "response_cache",
    "hit": True,
    "route": "/rag",
    "model": "model-x",
    "cache_key_version": "v3",
    "latency_ms": 18,
    "saved_latency_ms": 1240,
    "saved_input_tokens": 4200,
    "saved_output_tokens": 700,
    "llm_calls_avoided": 1,
    "embedding_calls_avoided": 0,
    "tool_calls_avoided": 0,
    "estimated_cost_saved": 0.031,
    "tenant": "tenant_123",
}

Send it to your telemetry pipeline.

9.2 Instrument the cache itself

Instead of scattering metrics everywhere, wrap the cache:

import time
 
 
class ObservableCache:
    def __init__(self, cache, metrics):
        self.cache = cache
        self.metrics = metrics
 
    def get(self, key):
        start = time.perf_counter()
        value = self.cache.get(key)
        latency_ms = (time.perf_counter() - start) * 1000
 
        self.metrics.record(cache="response", hit=value is not None, latency_ms=latency_ms)
        return value
 
    def set(self, key, value, ttl):
        start = time.perf_counter()
        self.cache.set(key, value, ttl)
        latency_ms = (time.perf_counter() - start) * 1000
 
        self.metrics.record_write(cache="response", latency_ms=latency_ms)

Now every cache operation is observable.

9.3 Track cache metrics by layer

A production AI system might stack several caches:

  1. Response cache
  2. Context cache
  3. Retrieval cache
  4. Embedding cache

Each layer needs its own metrics (response_hit_rate, context_hit_rate, retrieval_hit_rate, embedding_hit_rate). Don't collapse everything into one number.

9.4 The cache waterfall

For a RAG request, the path might be:

  1. Request
  2. Response cachemiss
  3. Query cachemiss
  4. Embedding cachehit
  5. Vector retrieval
  6. Reranker cachemiss
  7. Reranker
  8. Generation

Your observability should record the whole path:

trace = {
    "response_cache": "MISS",
    "query_cache": "MISS",
    "embedding_cache": "HIT",
    "retrieval_cache": "MISS",
    "reranker_cache": "MISS",
    "llm": "CALLED",
}

This makes bottlenecks obvious.

10. Cache economics

10.1 Find the most valuable cache

Suppose your metrics show:

Cache Hit rate Savings
Embedding 70% $100
Retrieval 40% $250
Reranker 60% $700
Response 30% $4,000

The response cache has the lowest hit rate, but it saves the most money.

Optimize based on value, not popularity.

10.2 Cost saved per cache hit

cost_saved_per_hit = total_cost_saved / total_cache_hits

For example, $4,000 saved over 20,000 hits is $0.20 per hit. Compare that with the lookup, storage and network cost of each hit.

10.3 Net savings

A better metric:

Net savings = avoided work cost − cache infrastructure cost − cache lookup cost

net_savings = avoided_compute_cost - cache_lookup_cost - storage_cost

This tells you whether the cache is actually economically useful.

10.4 When should you disable a cache?

Sometimes the answer is: remove the cache. Disable it when:

  • the hit rate is extremely low, lookups aren't free, and entries rarely repeat
  • hits frequently return stale data
  • it consumes significant memory without meaningful savings
  • its complexity creates more operational risk than value

Caching is not automatically good.

10.5 Cache policy tuning

Observability should drive policy. Suppose a 24-hour TTL gives a 30% hit rate and a 12% stale rate. Try a 2-hour TTL:

TTL Hit rate Stale rate
24 hours 30% 12%
2 hours 24% 2%

If quality improves substantially, the lower hit rate may be worth it. A TTL that's too short means poor reuse; a TTL that's too long means stale answers. The right value comes from workload data.

11. Freshness

11.1 Measure cache age

For AI caches, measure how old each served entry is:

import time
 
cache_age_seconds = time.time() - created_at

Aggregate the P50, P95 and P99 age and the maximum age, then compare them with your freshness requirements. For example:

Use case Freshness budget
News assistant 60 seconds
Internal documentation 1 hour
Static product documentation 24 hours

11.2 Freshness budgets

A useful production question:

How stale can this result be before it becomes unacceptable?

Define it explicitly:

FRESHNESS_BUDGET = {
    "news": 60,
    "inventory": 10,
    "documentation": 3600,
    "static_faq": 86400,
}
 
 
def is_fresh(cache_age, category):
    return cache_age <= FRESHNESS_BUDGET[category]

Now freshness is an explicit system requirement.

12. Dashboards and alerts

12.1 An observability dashboard

A production dashboard should answer:

Area Metrics
Performance Hit rate, lookup latency (P50 / P95 / P99), latency saved
AI compute LLM, embedding, reranker and tool calls avoided; tokens saved
Cost LLM, embedding and tool cost saved; total cost saved; cache infrastructure cost; net savings; ROI
Correctness Hit quality, stale hit rate, invalidation lag, version mismatches, groundedness
Infrastructure Memory utilization, eviction rate, hot keys, stampedes, cache errors

12.2 Alerts

Don't alert just because the hit rate drops below 80%; that might be normal. Alert on meaningful failures:

  • the stale hit rate rises above a threshold
  • P95 cache latency rises above a threshold
  • the eviction rate suddenly spikes
  • the cache error rate increases
  • cost savings suddenly collapse
  • a cache stampede is detected
  • the correctness evaluation score drops

12.3 An example production alert

Yesterday Today
Response cache hit rate 61% 59%
Cost saved $2,400 / day $2,300 / day
Hit quality 97% 88%

The hit rate and savings look fine; nothing alarming there. But hit quality dropped from 97% to 88%, and that's serious. The cache is still working mechanically, but it's becoming less trustworthy. This is why:

Correctness metrics matter more than the cache hit rate.

13. Models of cache value

13.1 The AI cache effectiveness model

The lifecycle of a useful cache hit:

  1. Cache hit
  2. Correct?
  3. Fresh?
  4. Work avoided?
  5. Cost saved?
  6. Worth it?

13.2 A complete cache telemetry model

A production cache event should ideally answer:

Question Answered by
Who? Tenant, application, route
What? Cache type, operation
Which version? Model, prompt, index, schema
Did it hit? Hit or miss
How fast? Lookup latency
How much work did it avoid? Tokens, calls, compute
How much money did it save? Estimated cost
Was it correct? Evaluation result
Was it fresh? Age, freshness status

That's far more useful than cache_hit = true.

13.3 A complete production architecture

A mature AI caching observability architecture can look like:

  1. User
  2. API gateway
  3. Cache layer
  4. Hitreturn the resultMissrun the AI pipeline
  5. EmbeddingRetrievalReranker
  6. LLM
  7. Result
  8. Cache
  9. Telemetry
  10. MetricsTracesLogs
  11. Dashboards
  12. Decisions

13.4 The production cache scorecard

For every important cache, keep a scorecard. For example, a RAG response cache:

Metric Value
Hit rate 67%
Hit quality 98%
Fresh hit rate 96%
P95 lookup latency 8 ms
P95 latency saved 1.4 s
LLM calls avoided 42,000 / day
Tokens saved 180M / day
Cost saved $3,200 / day
Infrastructure cost $180 / day
Net savings $3,020 / day
ROI 16.8x
Eviction rate 4%
Stampede rate 0.02%
Status Healthy

Now you can manage caching as an engineering system.

14. Before you ship

14.1 Production checklist

Before calling your cache production-ready:

Metrics

  • Hit rate
  • Miss rate
  • Lookup latency
  • Latency saved
  • Tokens saved
  • Calls avoided
  • Cost saved
  • Memory usage
  • Evictions

Correctness

  • Hit quality
  • Freshness
  • Staleness
  • Version mismatches
  • Invalidation lag
  • Groundedness
  • Authorization checks

Reliability

  • Stampede detection
  • Hot-key detection
  • Cache failure handling
  • Fallback path
  • Timeout handling
  • Circuit breaking

Economics

  • Infrastructure cost
  • Cost per cache hit
  • Net savings
  • ROI
  • Savings per GB
  • Savings per request

Evaluation

  • Shadow cache
  • Sampled hit evaluation
  • A/B testing
  • Quality regression detection
  • Freshness validation

14.2 The biggest observability mistakes

  1. Tracking only cache_hit_rate.
  2. Ignoring cache lookup latency.
  3. Not measuring tokens saved.
  4. Not measuring calls avoided.
  5. Not measuring stale results.
  6. Not evaluating semantic cache quality.
  7. Ignoring cache infrastructure cost.
  8. Caching everything.
  9. Not detecting cache stampedes.
  10. Optimizing hit rate instead of business value.

15. The full picture

15.1 The most important formula

Think of cache value as:

Cache value = useful work avoided − cache cost − correctness risk

where useful work includes LLM computation, embedding computation, retrieval, reranking, tool calls, database work, network calls and GPU compute. That is what your observability system should expose.

15.2 From cache hit rate to cache economics

Cache observability matures in levels:

  1. Did we hit?
  2. How fast was the hit?
  3. How much work did we avoid?
  4. How much money did we save?
  5. Was the cached result still correct?
  6. Was it fresh enough?
  7. Was the cache economically worth operating?

That's production-grade cache observability.

15.3 The AI caching measurement stack

  1. AI cache
  2. Performancelatency, hit rate, evictionsEconomicscost saved, tokens saved, calls avoidedCorrectnessfreshness, quality, groundedness
  3. ROI / value

The cache isn't successful because Redis says "HIT". It's successful when latency, cost, compute and API calls go down, while quality stays the same or better and freshness and reliability stay acceptable.

15.4 Part 9 takeaways

If you're building RAG, agentic RAG, conversational agents or other AI applications, remember:

  1. Hit rate is only the beginning. A 90% hit rate doesn't automatically mean a successful cache.
  2. Measure avoided work: LLM calls, embedding calls, tool calls, tokens and GPU compute.
  3. Measure money: cost saved − cache cost = net savings.
  4. Measure correctness. A stale cache hit is still a failure.
  5. Measure freshness. Different data needs different freshness budgets.
  6. Evaluate semantic caches. Semantic similarity doesn't guarantee equivalent answers.
  7. Use shadow caching. Test cache behaviour before trusting it.
  8. Monitor infrastructure: memory, evictions, hot keys and stampedes matter.
  9. Optimize for value. The best cache isn't necessarily the one with the highest hit rate. It's the one that prevents the most valuable work at acceptable correctness and infrastructure cost.

15.5 The production principle

Don't measure how often the cache is hit. Measure how much useful work the cache prevented.

That's the difference between "I added a cache" and "I built an economically measurable AI caching system".

15.6 The AI Caching Playbook so far

  1. Part 1: The 10 core caches
  2. Part 2: Agentic AI caching
  3. Part 3: Conversational AI caching
  4. Part 4: RAG caching
  5. Part 5: Cache the agent's work
  6. Part 6: LLM inference caching
  7. Part 7: Distributed AI caching
  8. Part 8: Cache invalidation and correctness
  9. Part 9: Cache observability and cost (this part)

What's next?

Part 10: The complete production AI caching architecture

We'll bring everything together:

  1. User
  2. API gateway
  3. Cache orchestrator
  4. Response cacheSemantic cacheSession cache
  5. AI workflow
  6. Embedding cacheRetrieval cacheTool cache
  7. RAG
  8. Agent
  9. LLM
  10. Prefix / KV cacheResponse cache
  11. Observability
  12. CostQualityLatency
  13. Optimization

The final question:

How do you design the entire caching layer as a first-class part of an AI system, not as an afterthought?

Next: Part 10, The complete production architecture brings every layer together into one production design.