HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 6: LLM Inference Caching

KV cache, prefix cache and prompt caching inside LLM inference: how they differ, what they cost in GPU memory, and how to design prompts for reuse.

Haribaskar Dhanabalan24 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

What if the model didn't have to recompute the same attention work every time?

This is Part 6 of the AI Caching Playbook. The earlier parts looked at caching around the model:

But there is another layer where enormous amounts of computation happen: inside LLM inference itself.

Every time an autoregressive LLM generates a token, it performs attention over the tokens that came before it. For a long prompt, recomputing all of that work for every generated token would be extremely expensive.

This is where the KV cache becomes one of the most important optimization techniques in modern LLM serving. Then there are related concepts, each one layer further down:

  1. Application cache
  2. Prompt cache
  3. Prefix cache
  4. KV cache
  5. GPU memory
  6. Inference engine

These are related, but they are not the same thing. Understanding the distinction is critical when building production AI systems.

1. The inference pipeline

1.1 Prefill and decode

Before discussing caching, understand what happens during generation. Suppose the user sends:

"Explain how RAG works in production."

The model receives tokens (Explain, how, RAG, works, in, production, .), processes the prompt, and then generates new tokens one at a time: RAG, systems, retrieve, relevant…

There are two major phases, and they behave very differently:

  1. LLM inference
  2. Prefillprocess the prompt, build the KV cacheDecodegenerate tokens, reuse the KV cache

1.2 The prefill phase

The prefill phase processes the input prompt: the system prompt, conversation history, retrieved documents and the user question. Conceptually:

prompt_tokens = tokenizer(prompt)
 
kv_cache = model.prefill(prompt_tokens)

The result is the internal attention state needed for generation. That state is made of K (keys) and V (values), hence the name: KV cache.

1.3 The decode phase

After prefill, the model starts generating tokens: "The", "answer", "is"…

Without a KV cache, the model would recompute attention information for all previous tokens at every step:

Generating Recomputes
Token 1 Everything before it
Token 2 Token 1 + token 2
Token 3 Tokens 1–3
Token 4 Tokens 1–4

With KV caching, K and V are computed once and reused at every step:

  1. Prompt
  2. Compute K/V once
  3. Store KV cache
  4. Token 1reuse KV
  5. Token 2reuse KV
  6. Token 3reuse KV

This dramatically reduces redundant computation during autoregressive generation.

1.4 What exactly is the KV cache?

Transformers use attention. At a simplified level, attention combines a query (Q), keys (K) and values (V): Attention(Q, K, V).

For tokens that have already been processed, the model's key and value tensors can be reused. Instead of recomputing K₁ K₂ K₃ K₄… and V₁ V₂ V₃ V₄… for every new token, the inference engine stores them, with a K and a V tensor for every layer of the model.

When a new token arrives:

  1. New query + existing K/V
  2. Attention
  3. Next token

This is one of the fundamental optimizations that makes autoregressive LLM generation practical.

2. KV cache is not prompt cache

2.1 Three different layers

This distinction is extremely important. People often use these terms interchangeably. They shouldn't.

Layer What it stores Scope
KV cache Internal attention tensors Usually an active inference request or session
Prompt cache Previously processed prompt or prefix computation Depends on the serving system or provider; may use KV state internally
Application cache Prompts, responses, embeddings, retrieval results, tool results Your application

They sit at different levels of the stack:

  1. Applicationresponse, RAG and tool caches
  2. Prompt / prefix reuse
  3. KV cache
  4. GPU

3. KV cache memory

3.1 KV memory can become huge

The biggest problem with KV caching is memory. The cache lives in accelerator memory during inference, and as context length increases, so does everything else:

  1. More tokens
  2. More K/V tensors
  3. More GPU memory

A rough memory calculation for standard attention:

KV memory ≈ batch × sequence_length × layers × 2 × KV_heads × head_dim × bytes_per_element

where the 2 is for K + V. In code:

def kv_cache_memory(batch, sequence_length, layers, kv_heads, head_dim, bytes_per_element):
    return batch * sequence_length * layers * 2 * kv_heads * head_dim * bytes_per_element
 
 
def bytes_to_gib(value):
    return value / (1024**3)
 
 
memory = kv_cache_memory(
    batch=1,
    sequence_length=8192,
    layers=32,
    kv_heads=8,
    head_dim=128,
    bytes_per_element=2,
)
 
print(bytes_to_gib(memory))  # 1.0 GiB for a single 8K-token sequence

Multiply that by longer contexts and many concurrent requests, and it's clear why long-context inference can become extremely memory-intensive.

3.2 MHA vs GQA vs MQA

KV memory depends heavily on the number of KV heads, not simply the total number of attention heads:

Attention type Heads KV memory
Multi-head attention (MHA) Q heads = K heads = V heads Large
Grouped-query attention (GQA) Many Q heads share fewer K/V heads Reduced
Multi-query attention (MQA) Many Q heads share one K/V head Even smaller

This is one reason modern model architectures are designed with efficient KV storage in mind.

3.3 KV cache and GQA

Suppose a model has 32 attention heads but only 8 KV heads. With GQA, Q has 32 heads while K and V have 8 each. That is a quarter of the KV storage of a model where K and V also have 32 heads.

The chain of consequences:

  1. Model architecture
  2. KV heads
  3. KV memory
  4. Concurrency
  5. Serving cost

Caching efficiency is partly decided before your application even starts.

4. Prefix caching

4.1 What is prefix caching?

Now we move from request-level KV caching to prefix reuse. Imagine every request starts with the same system prompt:

You are an AI assistant.
 
You must:
- answer accurately
- cite sources
- never invent facts
- follow company policy
...

Requests A, B and C each contain the same system prompt and the same developer instructions, followed by a different user message. The beginning is identical.

Instead of computing the shared prefix again for every request, a serving system can reuse the previously computed prefix state:

  1. Shared prefix
  2. User AUser BUser C

This is prefix caching, or prefix reuse.

4.2 Prefix matching is exact

This is an important production detail. Prefix caching generally depends on exact token-prefix reuse. It is not semantic caching.

"Explain RAG." and "Can you explain how retrieval augmented generation works?" are semantically similar, but their token sequences are different. A prefix cache can't decide they mean roughly the same thing and reuse the same KV state.

Cache Matches on
Semantic cache Similar meaning
Prefix cache An exact token prefix

Semantic similarity belongs to a different caching layer.

4.3 Design prompts for prefix reuse

This has a major architectural implication. If you want maximum prefix reuse, put stable content first.

  1. System prompt
  2. Current timestamp
  3. Random request ID
  4. User question
  5. Large static instructions
Bad: changing fields early in the prompt destroy prefix reuse
  1. Stable system instructions
  2. Stable tools
  3. Stable policies
  4. Stable examples
  5. Dynamic context
  6. User question
Better: stable content first, dynamic content last

For example:

Static prefix Dynamic suffix
System instructions User profile
Tool definitions Retrieved context
Output schema Conversation state
Safety policy Current question
Few-shot examples

This makes the stable portion reusable.

4.4 Prompt caching

At the application or provider level, you may hear the term prompt caching. The idea:

  1. Repeated prompt prefix
  2. Previously processed state
  3. Reuse
  4. Lower latency and compute

But don't assume that every API calling itself "prompt caching" works the same way. Different inference providers and serving engines implement it differently. Always distinguish between:

  • API-level prompt caching
  • inference-engine prefix caching
  • the raw KV cache

The architectural principle is the same, avoid recomputing identical prompt computation, but the implementation and guarantees are system-specific.

4.5 Cache key for prefix caching

If you build your own prefix-cache layer, the cache identity cannot simply be hash(prompt). A production key may need to account for:

  • the model and model version
  • the tokenizer version
  • the chat template version
  • the adapter / LoRA version
  • the exact token prefix
  • relevant inference configuration

For example:

import hashlib
 
 
def prefix_cache_key(
    model,
    model_version,
    tokenizer_version,
    template_version,
    prefix_tokens,
    adapter_version=None,
):
    payload = "|".join([
        model,
        model_version,
        tokenizer_version,
        template_version,
        adapter_version or "none",
        prefix_tokens,
    ])
    digest = hashlib.sha256(payload.encode()).hexdigest()
    return f"prefix:{digest}"

Why include the tokenizer and template versions? Because the same text does not necessarily mean the same tokenization, or the same model processing.

4.6 Chat template changes can break cache compatibility

Consider model v1 with chat template v1, which wraps the system prompt in <system>…</system>. Then you change to chat template v2. Even if the human-readable prompt looks similar, the actual token sequence can change.

  1. Prompt
  2. Tokenizer
  3. Chat template
  4. Tokens
  5. KV state

If any of these change, cached inference state may no longer be reusable. This is another example of why version-aware cache keys matter.

5. Reuse across turns and branches

5.1 Multi-turn KV reuse

Consider a conversation:

User: Explain RAG.

Assistant: RAG combines retrieval with generation…

User: What about reranking?

The second request contains the previous conversation.

  1. Entire conversation
  2. Process everything again
  3. Generate answer
Without reuse
  1. Previous conversation
  2. Existing KV
  3. New user message
  4. Generate
With session-level KV reuse

Conceptually:

session = {
    "kv_cache": existing_kv,
    "tokens": previous_tokens,
}
 
new_tokens = tokenizer(new_message)
 
output = model.decode(new_tokens, kv_cache=session["kv_cache"])

The exact implementation depends on the inference engine.

5.2 What happens when the user edits an earlier message?

This is where things get interesting. Suppose a conversation has five turns, A to E, and the user edits turn B.

The KV state computed after B is no longer valid for C, D and E, because those states were computed from the old history. Everything after the changed token may need to be recomputed:

  1. A
  2. B
  3. C
  4. D
  5. E
Original
  1. Astill valid
  2. B′
  3. Crecompute
  4. Drecompute
  5. Erecompute
After editing B

So conversational systems should treat the KV cache as dependent state, not as a permanent answer cache.

5.3 KV cache branching

This becomes useful for agentic systems. Suppose an agent explores two possible plans from the same context. Both branches can share the common prefix:

  1. Shared prefix
  2. Shared KV
  3. Plan AKV-APlan BKV-B

This is useful for beam search, tree search, agent planning, generating multiple candidates and reranking candidate responses.

5.4 Prefix sharing across requests

The bigger opportunity is cross-request reuse. Imagine a company chatbot where every request includes the company policy, product documentation, tool definitions and safety rules. That shared prefix could be huge.

Instead of computing the same 50K tokens for request 1, then again for request 2, then again for request 3, you want:

  1. Shared 50K-token prefix
  2. Cache
  3. Request 1Request 2Request 3

This is especially valuable for long system prompts, large tool definitions, agent policies, long documentation prefixes, repeated enterprise instructions and long structured contexts.

6. Prompt layout for reuse

6.1 Dynamic content should usually come later

Consider a RAG agent.

  1. System
  2. User query
  3. Retrieved documents
  4. Tool definitions
Bad: the user query changes early in every prompt
  1. System
  2. Tool definitions
  3. Policies
  4. Output schema
  5. Stable instructions
  6. Retrieved documents
  7. Conversation
  8. User query
Better: a much larger stable prefix

Conceptually, the prompt splits into two blocks:

Cacheable Dynamic
System User
Tools Retrieved docs
Policies Current state
Schemas Current question
Static examples

This is a powerful prompt-engineering technique for production inference.

6.2 Long-context caching

Long-context applications are where caching becomes particularly important. Suppose a 100K-token context contains system instructions, company documentation, previous conversation, retrieved documents and the user query. Processing the entire context every time can be expensive. Instead:

  1. 100K context
  2. Stable prefix
  3. Prefix cache
  4. Process only the new content

But caching does not make long context free. You still have memory cost, cache management, context limits, attention cost and eviction pressure.

Caching reduces repeated computation; it doesn't remove the underlying context.

7. Managing GPU memory

7.1 KV cache eviction

GPU memory is finite. Suppose the inference server has 64 GB of GPU memory and thousands of concurrent sessions. You cannot keep every KV cache forever, so you need eviction. Common strategies include:

  • LRU
  • TTL
  • LRU + TTL
  • priority-based eviction
  • session-aware eviction
  • size-aware eviction

A simplified cache:

from collections import OrderedDict
 
 
class KVCacheManager:
    def __init__(self, max_items=100):
        self.cache = OrderedDict()
        self.max_items = max_items
 
    def get(self, key):
        if key not in self.cache:
            return None
 
        value = self.cache.pop(key)
        self.cache[key] = value
        return value
 
    def put(self, key, value):
        if key in self.cache:
            self.cache.pop(key)
 
        self.cache[key] = value
 
        while len(self.cache) > self.max_items:
            self.cache.popitem(last=False)

Real inference engines need far more sophisticated memory management, because KV entries are large tensor blocks rather than ordinary Python objects. But the architectural principle is the same:

GPU memory is a scarce cache resource.

7.2 Paged KV cache

A major problem with naive KV allocation is fragmentation. Imagine allocating 10 GB for request A, 5 GB for request B and 15 GB for request C. When B finishes, it leaves a hole. Repeated allocation and deallocation creates an inefficient memory layout.

Paged KV-cache designs address this by managing KV memory in smaller blocks, or pages, instead of requiring each sequence to occupy one giant contiguous region:

  1. AABCCA
GPU memory in pages: each sequence's blocks can sit anywhere

This lets inference servers manage memory much more efficiently.

7.3 Why paged KV memory matters

  1. GPU memory
  2. Fragmentation
  3. Lower utilization
  4. Fewer concurrent requests
  5. Lower throughput
Without efficient memory management
  1. GPU memory
  2. Efficient allocation
  3. Higher utilization
  4. More concurrent sequences
  5. Higher throughput
With block-based management

This is particularly important for high-concurrency LLM serving.

7.4 KV cache quantization

The KV cache itself consumes a lot of memory. One strategy is to store it at lower precision:

  1. FP32
  2. FP16 / BF16
  3. Lower-precision formats

Conceptually:

kv_cache = kv_cache.to(dtype)

Lower precision can reduce memory, bandwidth and memory pressure. But it can introduce numerical error, quality degradation and stability issues. So:

Never assume that lower KV precision is automatically better. Benchmark it against your workload.

8. Scheduling and latency

8.1 Continuous batching

Caching is only one part of high-performance inference. Imagine requests A, B, C and D all generating at once. A naive server may process them inefficiently. Modern inference engines dynamically combine active sequences into batches:

  1. GPU
  2. Request A tokenRequest B tokenRequest C tokenRequest D token

As requests finish, new ones join: when A finishes, request E takes its place. This is commonly called continuous batching.

It is not itself a cache, but it interacts heavily with KV-cache memory.

8.2 KV cache and continuous batching

Continuous batching plus KV cache management is what gives efficient GPU utilization. The scheduler needs to know:

  • How much KV memory does each request need?
  • How much has already been allocated?
  • Which sequences are close to completion?
  • Which new request can fit?
  • Which cache blocks can be released?

So production inference is really a combination of a scheduler + batching + a KV memory manager + GPU execution.

8.3 TTFT vs token generation latency

When measuring inference performance, don't look only at total latency. Two important metrics:

Metric Measures Mainly influenced by
TTFT (time to first token) How long the user waits before the first generated token Prompt processing, prefill, queueing, prefix cache hit or miss
Inter-token latency How quickly the following tokens arrive Decode, KV cache, GPU utilization, batching

So:

  1. Prompt cache hit
  2. Less prefill work
  3. Lower TTFT

while:

  1. Efficient KV cache
  2. Efficient decode
  3. Faster token generation

9. Measuring prefix caching

9.1 Prefix cache hit rate

If you implement prefix caching, monitor prefix_cache_hits and prefix_cache_misses:

hit_rate = prefix_hits / max(prefix_hits + prefix_misses, 1)

But don't optimize hit rate blindly. A 95% hit rate is meaningless if the cached prefix is only 100 tokens, while a 50K-token prefix with a 10% hit rate could save dramatically more computation.

So track tokens reused, not just requests hit.

9.2 Track tokens reused

A better metric is cached tokens ÷ total prompt tokens. For example, if 70 million of 100 million prompt tokens were reused, 70% of prompt tokens came from the cache.

That tells you much more about the actual value of prefix caching.

9.3 Cache observability

A production inference system should expose metrics such as:

Area Metrics
KV cache Utilization, memory, evictions, allocation failures
Prefix cache Hit rate, miss rate, prefix tokens reused
Tokens Prompt tokens, cached prompt tokens, generated tokens
Latency Prefill latency, decode latency, TTFT, inter-token latency
GPU Utilization, memory utilization
Requests Queued, running, completed
Failures OOM events

For example:

metrics = {
    "kv_memory_gb": 42.3,
    "kv_utilization": 0.81,
    "prefix_hit_rate": 0.73,
    "cached_prompt_tokens": 18_400_000,
    "prompt_tokens": 25_000_000,
    "ttft_ms": 420,
    "decode_tokens_per_sec": 78,
}

10. Invalidation and versioning

10.1 Prompt cache invalidation

Just like every other cache, cached inference state can become stale or incompatible.

Imagine system prompt v1 is in the prefix cache, and then you deploy system prompt v2. You should not reuse the old prefix. Version it: system_prompt:v1 becomes system_prompt:v2, so the keys become prefix:v1:… and prefix:v2:…. Old entries naturally become unreachable.

This is much safer than trying to delete every old key immediately.

10.2 Version everything important

For production inference caching, consider versioning the model, tokenizer, chat template, system prompt, tool schema, adapter / LoRA, safety policy and inference runtime. For example:

cache_namespace = (
    f"{model_version}:"
    f"{tokenizer_version}:"
    f"{template_version}:"
    f"{prompt_version}:"
)

giving keys like prefix:model-v5:tokenizer-v3:template-v7:prompt-v12:<hash>. This prevents accidental reuse across incompatible versions.

10.3 Cache dependencies matter

A prefix is not isolated. It depends on the model, tokenizer, template, system prompt, tools, policy, schema and adapter. Think of it as a dependency graph:

  1. Model
  2. TokenizerChat templateAdapter
  3. System prompt
  4. Tool definitions
  5. Prefix
  6. KV state

If an upstream dependency changes, invalidate the downstream state. This is the same dependency-aware caching principle we saw in RAG and agent systems.

11. Other inference optimizations

11.1 Speculative decoding is different

Another optimization often mentioned alongside KV caching is speculative decoding. A smaller, faster model proposes tokens, and the larger model verifies them:

  1. Small model
  2. Generate candidate tokens
  3. Large model
  4. Verify
  5. Accept many tokens

This can improve generation efficiency. But:

Speculative decoding is not a cache. It is a decoding optimization.

Similarly, continuous batching is not a cache, and quantization is not a cache. They are complementary inference optimizations.

11.2 The inference optimization stack

A production LLM server may combine all of them:

  1. LLM serving
  2. Schedulingcontinuous batchingMemoryKV cache, prefix cache, paged memoryComputationquantization, speculative decoding

The goal: more tokens, lower latency, less GPU memory, higher concurrency and lower cost.

12. Cache-friendly architecture

12.1 Cache-friendly prompt architecture

For production applications, design the prompt intentionally:

  1. Stablesystem prompt, policies, tools, output schema, static examples, model instructions
  2. Prefix cache
  3. Dynamicuser profile, retrieved context, conversation state, current message
  4. Generation

This is one of the simplest changes an AI engineer can make to improve cacheability.

12.2 Cache-friendly RAG

Now combine this with RAG. A prompt made of system instructions, tool definitions, an output schema, retrieved documents, conversation and the question can be arranged as:

Static prefix Dynamic context Dynamic query
System, tools, policies, schema, examples Retrieved documents, conversation User question

Then:

  1. Prefix cache
  2. RAG retrieval
  3. Dynamic context
  4. Generation

Notice the distinction: a RAG cache reuses retrieval results, while a prefix cache reuses model computation. The two layers work together.

12.3 Cache-friendly agent architecture

For an agent, the same split applies:

  1. Stablesystem, agent policy, tools, schemas, safety rules, output format
  2. Prefix cache
  3. Dynamicmemory, task state, tool outputs, current user input
  4. LLM

This can significantly speed up repeated agent execution when the stable portion is large.

12.4 A practical prefix builder

A simple application-level abstraction:

class PromptBuilder:
    def __init__(self, system_prompt, tools, policy, schema):
        self.system_prompt = system_prompt
        self.tools = tools
        self.policy = policy
        self.schema = schema
 
    def build_static_prefix(self):
        return {
            "system": self.system_prompt,
            "tools": self.tools,
            "policy": self.policy,
            "schema": self.schema,
        }
 
    def build_dynamic_context(self, memory, retrieved_docs, user_message):
        return {
            "memory": memory,
            "retrieved_docs": retrieved_docs,
            "user_message": user_message,
        }

Then:

prefix = builder.build_static_prefix()
 
dynamic = builder.build_dynamic_context(
    memory=memory,
    retrieved_docs=docs,
    user_message=user_message,
)

The architecture makes the separation explicit: the static part is cacheable, the dynamic part is request-specific.

12.5 A prefix fingerprint

You can also fingerprint the stable portion:

import hashlib
import json
 
 
def fingerprint(value):
    payload = json.dumps(value, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(payload.encode()).hexdigest()
 
 
prefix = builder.build_static_prefix()
prefix_id = fingerprint(prefix)

Now prefix_id identifies the prefix cache entry. If the system prompt, tools, policy or schema changes, prefix_id changes, and old cache entries are naturally bypassed.

13. Security

13.1 Inference caches hold sensitive data

Inference caches can contain highly sensitive information. A KV cache may indirectly represent user conversations, private documents, company policies, personal information, tool outputs and retrieved data.

So never assume a cache is harmless temporary data. In multi-tenant systems, each tenant gets its own cache; never share state across tenants by accident. A safe cache namespace might include:

key = (
    f"tenant:{tenant_id}:"
    f"model:{model_version}:"
    f"prefix:{prefix_hash}"
)

Even then, authorization must still be enforced independently.

13.2 Never use the cache as authorization

This is a critical production rule. Bad: the cache says user X can access document Y, so access is allowed. Don't do this. Instead:

  1. Authentication
  2. Authorization
  3. Cache lookup
  4. Use cached data if authorized

Caching can speed up a decision. It should not replace your security boundary.

13.3 Cache isolation

Depending on your application, you may need isolation by tenant, user, session, conversation, model, model version, prompt version and policy version. For example:

tenant:acme / user:123 / session:abc / model:xyz / prompt:v8 / prefix:<hash>

The exact dimensions depend on whether the cached state is actually shareable.

14. Putting it together

14.1 The complete production flow

A production request can look like this:

  1. User request
  2. API gateway
  3. Authentication
  4. Authorization
  5. Prompt builder
  6. Static prefixDynamic data
  7. Prefix cacheRAG / memory
  8. Scheduler
  9. Continuous batching
  10. KV manager
  11. GPU
  12. Decode
  13. Response

This is where all the caching concepts start connecting.

14.2 The complete cache stack

At this point, an AI system can have caching at many layers, each solving a different problem:

  1. Response cache
  2. Conversation cache
  3. Agent cache
  4. RAG cache
  5. Embedding cache
  6. Retrieval cache
  7. Prompt / prefix cache
  8. KV cache
  9. GPU memory

14.3 Don't confuse the cache layers

A useful mental model:

Cache What it stores Main benefit
Response cache Final answer Avoid generation
Semantic cache Similar Q&A Avoid repeated requests
RAG cache Retrieval results Avoid repeated retrieval
Embedding cache Vectors Avoid embedding computation
Tool cache Tool outputs Avoid external calls
Prompt cache Reusable prompt processing Reduce prompt compute
Prefix cache Reusable prompt prefix state Reduce prefill work
KV cache Attention K/V tensors Efficient decoding
GPU memory Runtime state Fast inference

The important part:

These caches are complementary, not interchangeable.

14.4 Cost optimization through layering

When a user asks a question, you can potentially avoid work at several layers:

Level Cache When it applies What you avoid
1 Exact response cache Same request The model call entirely
2 Semantic cache Similar request, validated response Potentially the model call
3 RAG cache Same query Retrieval work (the model is still called)
4 Embedding cache Same text Embedding computation
5 Prefix cache Same prompt prefix Prefill work
6 KV cache Same active sequence Repeated decode computation

This gives you a layered optimization strategy.

15. Deciding what to optimize

15.1 Measure before optimizing

Don't add every cache blindly. First measure:

  • Where is the latency?
  • Where is the cost?
  • Where is GPU memory going?
  • Which requests repeat?
  • Which prefixes repeat?
  • How large are the repeated prefixes?
  • How often does context change?

For example:

Stage Time
RAG retrieval 80 ms
Prompt prefill 900 ms
Tool call 1.5 s
Generation 2.2 s

Your priority should follow the numbers. If prompt prefill dominates, prefix caching may be valuable. If tool calls dominate, tool-result caching may matter more. If retrieval dominates, RAG caching may give a better return.

15.2 Cache hit rate isn't the goal

This deserves repeating. The goal is not a 100% cache hit rate. The goal is lower cost, lower latency and higher throughput, with no correctness regression.

Hit rate Savings
Cache A 90% $5 / month
Cache B 30% $20,000 / month

Cache B is far more valuable. So measure latency saved, tokens avoided, GPU compute avoided, GPU memory efficiency and cost saved.

15.3 A production mental model

When debugging an LLM application, ask these questions in order:

  1. Can I avoid the request entirely?
  2. Can I reuse the response?
  3. Can I reuse retrieval?
  4. Can I reuse embeddings?
  5. Can I reuse the prompt prefix?
  6. Can I reuse KV state?
  7. Can I make decoding more efficient?
  8. Can I improve GPU utilization?

This gives you a complete inference optimization ladder.

16. Before you ship

16.1 Production checklist

Before deploying inference caching, ask:

KV cache

  • Are KV tensors reused during decoding?
  • How much GPU memory does each sequence consume?
  • What is the maximum concurrent context?
  • What happens when memory is exhausted?
  • How are KV blocks allocated?
  • How are KV entries evicted?

Prefix cache

  • Which prompt sections are stable?
  • Are prefixes token-identical?
  • Are dynamic fields placed after stable content?
  • Is prefix reuse measured?
  • How many tokens are reused?

Versioning

  • Model version?
  • Tokenizer version?
  • Chat template version?
  • Prompt version?
  • Tool schema version?
  • Adapter version?

Memory

  • KV cache utilization?
  • GPU memory utilization?
  • Fragmentation?
  • Evictions?
  • OOM rate?

Performance

  • TTFT?
  • Prefill latency?
  • Decode latency?
  • Tokens per second?
  • Queue time?
  • Cache hit rate?
  • Tokens reused?

Security

  • Tenant isolation?
  • User and session isolation?
  • Sensitive data handling?
  • Authorization independent of the cache?

16.2 A production metrics dashboard

A useful dashboard could look like:

Area Metric Value
LLM inference Requests / sec 142
LLM inference Tokens / sec 9,842
LLM inference TTFT 310 ms
LLM inference P95 TTFT 720 ms
LLM inference Decode speed 81 tok/s
LLM inference P95 latency 3.8 s
KV cache GPU KV memory 71%
KV cache KV eviction rate 4.2%
KV cache KV allocation failures 0.03%
Prefix cache Hit rate 68%
Prefix cache Tokens reused 72%
Prefix cache Cache misses 32%
Prefix cache Prompt tokens 41.2M
Prefix cache Cached tokens 29.7M
GPU Utilization 88%
GPU Memory 91%
GPU OOM events 0

This gives you a much better picture than monitoring API latency alone.

17. The full picture

17.1 The bigger architecture

Now combine everything from across the series:

  1. User
  2. API gateway
  3. Auth / ACL
  4. Response cache
  5. Conversation cache
  6. Agent cache
  7. RAG cache
  8. Embedding cacheRetrieval cacheReranker cache
  9. Context
  10. Prompt builder
  11. Stable prefixDynamic
  12. Prefix cacheMemory / RAG
  13. Scheduler
  14. Continuous batching
  15. KV manager
  16. GPU
  17. PrefillDecodereuse KV
  18. Response

This is the real production picture. Caching isn't one feature. It is an architecture spanning the entire AI stack.

17.2 The core principle

The biggest lesson from this part:

Don't make the model recompute work you already paid for.

Level Cache
Application The answer
RAG The retrieval
Embedding The vector
Agent The work
Prompt The reusable prefix
Inference The attention state
Decoding Reuse the KV cache

17.3 The AI Caching Playbook so far

Across the series:

  1. Part 1: The 10 core caches
  2. Part 2: Agentic AI caching
  3. Part 3: Conversational AI caching
  4. Part 4: RAG caching
  5. Part 5: Cache the agent's work
  6. Part 6: LLM inference caching (this part)

And the caching hierarchy becomes:

  1. AI system
  2. Applicationresponse, session, memoryRAGretrieval, embedding, rerankingAgentstools, state, workflow
  3. Prompt / prefix
  4. KV cache
  5. GPU

The deeper you go into the stack, the more expensive the cached state can become, and the more carefully it must be managed.

17.4 Part 6 takeaways

If you remember only these points:

  1. The KV cache is fundamental to autoregressive generation. It stores previously computed key/value attention states so the model doesn't recompute them for every generated token.
  2. The KV cache is not the same as a prompt cache. Prompt and prefix caching reuse previously processed prompt prefixes; the KV cache is the underlying inference state used during generation.
  3. Prefix reuse needs an exact token-prefix match. Semantic similarity is not enough.
  4. Put stable content first. Stable content is cacheable; dynamic content is request-specific.
  5. The KV cache consumes GPU memory. Long contexts and high concurrency can make KV memory one of the biggest serving constraints.
  6. GQA and MQA reduce KV memory. Fewer KV heads mean less memory pressure.
  7. Paged, block-based KV management improves utilization. Efficient memory allocation becomes critical at scale.
  8. Continuous batching and caching work together. Batching improves GPU utilization while the KV cache avoids redundant attention computation.
  9. Version your cache dependencies. At minimum: model, tokenizer, chat template, prompt, tools and adapters.
  10. Optimize for useful work avoided. Don't obsess over cache hit rate. Measure tokens reused, latency saved, GPU memory saved, cost saved and throughput gained.

17.5 The final mental model

When building a production AI system, think about caching from top to bottom:

  1. User request
  2. Response cache
  3. Conversation cache
  4. Agent cache
  5. RAG cache
  6. Embedding cache
  7. Retrieval cache
  8. Reranking cache
  9. Prompt cache
  10. Prefix cache
  11. KV cache
  12. GPU memory
  13. LLM engine

Every layer asks the same question:

What work have we already done that we don't need to do again?

That's the real idea behind caching in AI systems. Not "put Redis everywhere", but:

  • avoid redundant computation
  • avoid redundant retrieval
  • avoid redundant network calls
  • avoid redundant token processing
  • avoid redundant attention computation
  • reuse expensive state safely

Once you start looking at AI systems this way, caching stops being a performance afterthought. It becomes part of the core system architecture.

What's next?

Part 7: Distributed AI caching and production cache architecture

The next level is no longer just "What should I cache?". It becomes:

"How do I operate all these caches at production scale?"

We'll get into:

  • distributed caching, Redis Cluster, consistent hashing and cache sharding
  • multi-region caching, local + distributed caches, and L1/L2/L3 cache architecture
  • cache replication and cache invalidation at scale
  • distributed locks, cache stampede prevention and request coalescing
  • hot keys, cache warming, eviction strategies and TTL design
  • versioned namespaces and tenant isolation
  • failure handling, Redis memory management and observability
  • cost optimization, disaster recovery and production cache architecture

The question becomes:

How do you build a caching layer that stays fast, correct and reliable when your AI system handles millions of requests?

That is where caching becomes a distributed-systems problem.

Next: Part 7, Distributed AI caching covers sharding, hot keys, stampedes, invalidation at scale and failure handling.