HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 1: The 10 Core Caches

KV, prompt, semantic, embedding, retrieval, reranking, tool, API, conversation and agent state caches: what each saves and how to key it safely.

Haribaskar Dhanabalan15 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

Modern AI applications are rarely limited by model intelligence alone.

In production, the real problems are often:

  • Why is this request taking 8 seconds?
  • Why did our LLM bill double?
  • Why are we computing the same embedding thousands of times?
  • Why is the agent calling the same API repeatedly?
  • Why are users asking the same question, but every request still hits the model?
  • Why does our RAG pipeline perform the same retrieval work again and again?
  • Why does a long conversation become increasingly expensive?

The answer to many of these problems is caching.

Caching in AI is not one single technology. There are multiple cache layers operating at different points in the AI stack. A production RAG or agentic AI system may look like this:

  1. User
  2. Request cache
  3. Semantic cache
  4. Prompt cache
  5. LLM inference
  6. RetrievalTools
  7. Retrieval cacheTool result cache
  8. Reranker cacheAPI response cache
  9. Final answer
Cache layers around a single request

The important idea is:

Don't cache only the final answer. Cache expensive intermediate work too.

This article covers the first 10 cache layers every AI engineer should understand, in the order they sit in the stack:

  1. KV cache
  2. Prompt cache
  3. Semantic cache
  4. Embedding cache
  5. Retrieval cache
  6. Reranking cache
  7. Tool result cache
  8. API response cache
  9. Conversation cache
  10. Agent state cache

1. KV cache

1.1 What is it?

KV cache stands for key-value cache. It is one of the most important caches inside transformer inference.

When an LLM processes tokens, the attention mechanism creates keys and values for each one. For tokens that have already been processed, those keys and values don't need to be recomputed from scratch during autoregressive generation. The model can reuse them. Conceptually:

  1. Tokens 1…n
  2. Attention
  3. KV cache
  4. Next token

Without KV caching, every new token means recomputing the whole previous context:

  • Generate token 1 → recompute previous context
  • Generate token 2 → recompute previous context
  • Generate token 3 → recompute previous context
  • …and so on

With KV caching, the previous context is processed once and reused:

  1. Process previous context once
  2. KV cache
  3. Reuse during generation

This dramatically improves autoregressive decoding efficiency.

1.2 Where does the KV cache live?

Typically in GPU memory, next to the model weights and the activations. That makes the KV cache fundamentally an inference-engine and model-serving concern.

You generally don't implement KV caching yourself when using a managed LLM API. Inference engines such as vLLM, TensorRT-LLM and other serving systems manage it.

1.3 Why should RAG engineers care?

Because RAG increases context. Every request carries:

  1. User query
  2. System prompt
  3. Retrieved documents
  4. Conversation history

The larger the context, the more important efficient attention and KV management become. For long-context applications running at 10K, 50K or 100K+ tokens, KV memory can become a major infrastructure constraint.

1.4 Production considerations

Watch:

  • KV cache memory usage
  • Context length
  • Concurrent requests
  • GPU memory
  • Batch size
  • Prefix reuse
  • Cache eviction
  • Quantization
  • Paged KV cache

A production system should treat KV memory as a capacity-planning problem.

2. Prompt caching

2.1 What is it?

Prompt caching is different from KV caching. The idea is:

If a large portion of a prompt stays unchanged, don't repeatedly process the same prefix.

Consider an agent whose prompt contains system instructions, company policies, tool definitions, a large knowledge context and, at the end, the user's question. Only the user's question changes. Every request could look like:

  1. 100K tokens of static context
  2. 20 tokens of new user input

Processing that static prefix again and again is wasteful. Prompt caching lets the provider or inference system reuse the previously processed prefix, where supported.

2.2 Example

Request Without caching With prompt caching
1 Process 100K static tokens + 20 new tokens Process 100K static tokens (cached) + 20 new tokens
2 Process 100K static tokens again + 20 new tokens Reuse cached prefix + 20 new tokens
3 Process 100K static tokens again + 20 new tokens Reuse cached prefix + 20 new tokens

2.3 Ideal candidates

Prompt caching works particularly well for:

  • Large system prompts. "You are an enterprise support agent…" followed by a long policy document and detailed instructions.
  • Agent tool definitions. Tool 1, tool 2, tool 3… tool 50: the same schemas on every call.
  • Stable enterprise context. Company policies, product documentation, security rules, compliance instructions.

2.4 Production rule

Put stable content before dynamic content when your model or provider's caching benefits from prefix reuse.

  1. System prompt
  2. Company policy
  3. Tool definitions
  4. Knowledge
  5. User query
Cache-friendly: stable prefix, dynamic tail
  1. User query
  2. System prompt
  3. Random context
  4. Tool definitions
Less cache-friendly: the prefix changes every request

3. Semantic cache

3.1 What is it?

This is one of the most useful caches for RAG applications.

Traditional caching asks:

"Is the request exactly the same?"

Semantic caching asks:

"Is this request meaningfully similar to something we've already answered?"

Consider three questions:

  1. "What is the refund policy?"
  2. "Can I get my money back?"
  3. "How does your refund process work?"

They aren't identical strings, but they may have the same intent. A normal cache treats "What is the refund policy?" and "Can I get my money back?" as two different keys. A semantic cache embeds the question, runs a similarity search and finds the existing answer:

  1. "What is the refund policy?"
  2. Embedding
  3. Similarity search
  4. Existing answer

3.2 Basic architecture

  1. User query
  2. Generate embedding
  3. Search cache
  4. Similar enoughreturn cached responseNo matchcontinue
  5. RAG / LLM
  6. Store response

3.3 Python example

A simplified implementation can use an embedding model and a list (in production, a vector database):

from sentence_transformers import SentenceTransformer
import numpy as np
 
model = SentenceTransformer("all-MiniLM-L6-v2")
 
cache = []
 
 
def similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
 
 
def semantic_cache_lookup(query, threshold=0.90):
    query_embedding = model.encode(query)
 
    for item in cache:
        score = similarity(query_embedding, item["embedding"])
        if score >= threshold:
            return item["response"]
 
    return None

Store a response:

def store_cache(query, response):
    embedding = model.encode(query)
 
    cache.append({
        "query": query,
        "embedding": embedding,
        "response": response,
    })

Then:

query = "Can I get a refund?"
 
cached = semantic_cache_lookup(query)
 
if cached:
    answer = cached
else:
    answer = call_llm(query)
    store_cache(query, answer)

3.4 The dangerous part

Semantic cache thresholds are extremely important.

  • Too low (threshold = 0.70): you may return incorrect answers.
  • Too high (threshold = 0.99): you barely get cache hits.

Production systems should tune thresholds using real traffic.

4. Embedding cache

4.1 What is it?

Embedding generation is another place where repeated computation happens.

Suppose you ingest document.pdf and split it into 10,000 chunks, then generate an embedding for each chunk. If the same chunk is processed again, generating its embedding again is unnecessary.

4.2 Architecture

  1. Text
  2. Hashtext + embedding model + version
  3. Embedding cache
  4. Hitreturn embeddingMisscall embedding model
  5. Store in cache

4.3 Simple Python implementation

import hashlib
 
embedding_cache = {}
 
 
def cache_key(text, model_name):
    value = f"{model_name}:{text}"
    return hashlib.sha256(value.encode()).hexdigest()
 
 
def get_embedding(text, model_name, embed_fn):
    key = cache_key(text, model_name)
 
    if key in embedding_cache:
        return embedding_cache[key]
 
    embedding = embed_fn(text)
    embedding_cache[key] = embedding
    return embedding

4.4 Important production detail

The cache key should not be hash(text) alone. It should usually incorporate:

  • the text
  • the embedding model
  • the embedding model version
  • the preprocessing version

Otherwise:

  1. Old embedding model
  2. Old cached vector
  3. New embedding model
  4. Inconsistent vector space

That can silently damage retrieval quality.

5. Retrieval and vector search cache

5.1 What is it?

RAG systems repeatedly receive similar queries. "What is our leave policy?" might be asked hundreds of times. If your retrieval pipeline runs the full path every time, you're repeating work:

  1. Query
  2. Embedding
  3. Vector DB
  4. Top-k documents

A retrieval cache stores the result.

5.2 Architecture

  1. Query
  2. Normalize
  3. Cache key
  4. Retrieval cache
  5. Hitreturn documentsMissquery the vector database
  6. Top-k documents
  7. Store in cache

5.3 Example

retrieval_cache = {}
 
 
def retrieve(query, vector_db, top_k=10):
    key = f"{query}:{top_k}"
 
    if key in retrieval_cache:
        return retrieval_cache[key]
 
    results = vector_db.search(query=query, top_k=top_k)
    retrieval_cache[key] = results
    return results

5.4 But there is a major problem

Documents change.

On Monday, the refund policy says 30 days. You cache "What is the refund policy?" → 30 days. On Tuesday, the policy changes to 14 days. Your cache might still return 30 days.

This is a cache invalidation problem.

5.5 Production solution

Make these part of the cache key:

  • the document version
  • the index version
  • the query
  • the retrieval configuration

For example:

key = (
    f"{query}:"
    f"{index_version}:"
    f"{top_k}:"
    f"{retrieval_version}"
)

Now when the index changes, old entries simply stop matching:

  1. index_v1invalidated
  2. index_v2new retrieval

6. Reranking cache

6.1 What is it?

Modern RAG systems often use a two-stage pipeline:

  1. Retriever
  2. Top 50 documents
  3. Reranker
  4. Top 5 documents

Reranking can be expensive. For a repeated query with the same candidate documents, the reranking result can be cached.

6.2 Architecture

  1. Query + candidate IDs
  2. Reranker cache
  3. Hitreturn ranked resultsMissrun the reranker
  4. Store in cache

6.3 Cache key

Don't key on the query alone, because the candidate documents might change. Better:

import hashlib
 
 
def rerank_key(query, document_ids, model_version):
    value = query + ":" + ",".join(document_ids) + ":" + model_version
    return hashlib.sha256(value.encode()).hexdigest()

This protects against returning a ranking generated from a different document set.

7. Tool result cache

7.1 What is it?

This becomes extremely important in agentic AI.

Imagine an agent with a weather tool, a database tool, a CRM tool, a search tool, a pricing tool and an inventory tool. The agent might repeatedly call get_customer(123) or get_inventory("SKU-123").

If the underlying information doesn't change frequently, caching the result can save:

  • latency
  • API costs
  • database load
  • rate limits

7.2 Example

tool_cache = {}
 
 
def cached_tool_call(tool_name, arguments, tool_fn):
    key = (tool_name, tuple(sorted(arguments.items())))
 
    if key in tool_cache:
        return tool_cache[key]
 
    result = tool_fn(**arguments)
    tool_cache[key] = result
    return result

Usage:

customer = cached_tool_call(
    "get_customer",
    {"customer_id": "123"},
    get_customer,
)

7.3 The most important concept: TTL

Not every tool should have the same time to live (TTL). For example:

Tool TTL
Inventory 10 seconds
Exchange rate 1 minute
Weather 5 minutes
Customer profile 10 minutes
Product details 1 hour
Company policy 24 hours

Caching isn't simply cache = true. A cache policy is TTL + freshness + consistency + invalidation.

8. API response cache

8.1 What is it?

AI applications often depend on external APIs: weather, payments, CRM, search and internal services, all called by the LLM or the code around it. Repeated API calls can become expensive or slow. A response cache stores the API response.

8.2 Example

import time
 
api_cache = {}
 
 
def cached_api_call(key, api_fn, ttl=300):
    now = time.time()
 
    if key in api_cache:
        value, timestamp = api_cache[key]
        if now - timestamp < ttl:
            return value
 
    result = api_fn()
    api_cache[key] = (result, now)
    return result

Usage:

weather = cached_api_call(
    "weather:bengaluru",
    lambda: fetch_weather("bengaluru"),
    ttl=300,
)

8.3 Production consideration

Never blindly cache every API. Ask:

  • Is the response deterministic?
  • Can it become stale?
  • Does it contain private information?
  • Does the API prohibit caching?
  • What is the correct TTL?
  • Does the user have permission to see it?

9. Conversation and context cache

9.1 What is it?

Conversational agents repeatedly need access to conversation state. Imagine:

User: My order number is 12345.

Assistant: Got it.

User: Where is it now?

The second request needs context from the first. Instead of reconstructing the entire conversation from the database every time, applications often keep a session or context cache.

9.2 Architecture

The conversation ID is the key. The context cache holds:

  • recent messages
  • user preferences
  • the current task
  • tool results
  • session metadata

9.3 Redis example

import json
 
import redis
 
redis_client = redis.Redis(host="localhost", port=6379, decode_responses=True)
 
 
def save_conversation(conversation_id, messages):
    redis_client.setex(
        f"conversation:{conversation_id}",
        3600,
        json.dumps(messages),
    )
 
 
def get_conversation(conversation_id):
    data = redis_client.get(f"conversation:{conversation_id}")
 
    if not data:
        return []
 
    return json.loads(data)

9.4 Don't cache the entire conversation forever

Long conversations create another problem:

  1. 10 messages
  2. 50 messages
  3. 500 messages
  4. 5,000 messages

Eventually the context becomes expensive. Instead of sending the entire history every time, production systems often combine:

  • recent messages
  • a conversation summary
  • important facts
  • long-term memory

10. Agent state cache

10.1 What is it?

This is particularly important for agentic RAG. An agent isn't simply input → output. It may run:

  1. Goal
  2. Plan
  3. Step 1tool → result
  4. Step 2tool → result
  5. Decision
  6. Final answer

This intermediate state can be cached.

10.2 Example

Suppose an agent is researching a company. Its task has five steps:

  1. Search for the company
  2. Find the financial report
  3. Extract revenue
  4. Find competitors
  5. Generate a summary

If the agent crashes after step 4, you don't want to start again from step 1. Store its state:

{
  "task_id": "task_123",
  "status": "running",
  "completed_steps": [
    "company_search",
    "financial_report",
    "revenue_extraction",
    "competitor_search"
  ],
  "results": {
    "revenue": "...",
    "competitors": ["...", "..."]
  }
}

Now the agent can resume.

10.3 Simple implementation

agent_state = {}
 
 
def save_state(task_id, state):
    agent_state[task_id] = state
 
 
def load_state(task_id):
    return agent_state.get(task_id)

In the agent:

state = load_state(task_id)
 
if not state:
    state = {"completed_steps": [], "results": {}}
 
if "search_company" not in state["completed_steps"]:
    result = search_company()
 
    state["results"]["company"] = result
    state["completed_steps"].append("search_company")
 
    save_state(task_id, state)

This turns an agent from stateless execution into recoverable execution.

11. Putting the 10 caches together

A production RAG + agent architecture might look like this:

  1. User
  2. API / gateway
  3. Semantic cache
  4. Conversation context cache
  5. Agent
  6. RAGToolsAgent state cache
  7. Embedding cacheTool result cache
  8. Retrieval cacheRAG path continues
  9. Reranking cache
  10. LLM
  11. Prompt / KV cache

This is why "AI caching" is much bigger than caching the final LLM response.

12. Designing the cache layer

12.1 Cache keys are a first-class design problem

One of the biggest mistakes in production AI systems is a weak cache key.

Bad:

key = query

Better: include every version the result depends on.

import hashlib
 
 
def make_cache_key(
    query,
    model_version,
    prompt_version,
    knowledge_version,
    retrieval_version,
):
    raw = "|".join([
        query,
        model_version,
        prompt_version,
        knowledge_version,
        retrieval_version,
    ])
    return hashlib.sha256(raw.encode()).hexdigest()

This prevents old results from leaking into new system versions.

12.2 Cache invalidation is harder in AI

Traditional software already has the famous problem:

"There are only two hard things in Computer Science: cache invalidation and naming things."

AI systems make invalidation even more complicated. One change ripples through the whole stack:

  1. Document changes
  2. Embedding changes
  3. Vector index changes
  4. Retrieval changes
  5. Reranking changes
  6. Prompt changes
  7. Model changes

Your cached answer may depend on all of them. So production cache keys should often encode versions:

  • query
  • embedding_version
  • index_version
  • reranker_version
  • prompt_version
  • model_version

12.3 TTL strategy

Don't use one TTL for everything. A better starting point:

Cache Typical strategy
KV cache Inference lifecycle
Prompt cache Provider / inference policy
Semantic cache Minutes → hours
Embedding cache Long-lived
Retrieval cache Minutes → hours
Reranking cache Minutes → hours
Tool result cache Seconds → hours
API response cache API-dependent
Conversation cache Session-based
Agent state cache Until task completion

These are starting points, not universal values. Your actual TTL should be determined by data volatility, correctness requirements, cost, traffic, latency and compliance.

12.4 The security problem

Caching AI results can introduce serious security vulnerabilities. Consider:

  1. User A"What is my salary?"
  2. Response cached
  3. User Basks a similar question
  4. Cache hit
  5. User A's salary is returned to User B

That's a catastrophic failure. So cache keys may need:

  • tenant_id
  • user_id
  • permissions
  • the query
  • the version

For example:

key = (
    f"tenant:{tenant_id}:"
    f"user:{user_id}:"
    f"query:{query_hash}"
)

For multi-tenant AI applications:

Never assume a cache is safe to share across users or tenants.

12.5 Cache observability

If you introduce caching without observability, you won't know whether it actually helps. Track:

  • cache hit rate and miss rate
  • hit latency and miss latency
  • saved tokens, saved API calls and saved cost
  • stale response rate
  • invalidation rate and eviction rate
  • cache size
  • error rate

A useful dashboard might show:

Metric Value
Semantic cache hit rate 68%
Embedding cache hit rate 94%
Retrieval cache hit rate 41%
Tool cache hit rate 73%
Tokens saved 18.4M
API calls avoided 42,891
Estimated cost saved $1,284
Stale responses 0.03%
Cache errors 0.01%

The key metric isn't:

"How many cache hits do we have?"

It's:

"How much useful work did the cache eliminate without hurting correctness?"

12.6 Cache hit rate isn't everything

A 95% cache hit rate sounds great. But if the 5% of misses are the expensive requests, with 10-second latency, huge LLM calls and multiple tool calls, your system can still feel slow.

Track the expected values instead:

Expected latency = hit rate × hit latency + miss rate × miss latency
Expected cost    = hit rate × cache cost  + miss rate × full pipeline cost

12.7 A production cache policy

Instead of:

if cache:
    return cache

think:

def get_cached_result(key, ttl, tenant, permissions, version):
    # 1. Validate tenant isolation
    # 2. Validate permissions
    # 3. Validate version
    # 4. Check TTL
    # 5. Check freshness
    # 6. Return cached result
    # 7. Otherwise execute pipeline
    pass

Caching is not just a performance feature. It becomes part of your correctness architecture.

13. Deciding what to cache

13.1 A practical decision framework

Before adding a cache, ask five questions.

1. Is the computation expensive?

LLM inference, embeddings, reranking, external APIs and database queries usually are. If the work is cheap, caching may not matter.

2. Is the result reusable?

If every request is unique, the cache hit rate will be close to 0%. Don't cache it.

3. Can the result become stale?

If yes, you must design TTL + versioning + invalidation.

4. Can one user see another user's result?

If yes, you need tenant isolation + authorization-aware keys.

5. Can an incorrect cached result cause harm?

In healthcare, finance, legal, security and enterprise permissions, the answer may be yes. In these cases, cache aggressively only where correctness remains provable.

13.2 What should you cache first?

Don't implement all 10 at once. Start where your system repeats the most work.

  1. Embedding cache
  2. Retrieval cache
  3. Reranking cache
  4. Semantic cache
Typical RAG system
  1. Tool result cache
  2. Agent state cache
  3. Semantic cache
  4. Retrieval cache
Agentic RAG system
  1. Conversation / context cache
  2. Prompt cache
  3. Semantic cache
  4. Tool result cache
Conversational agent
  1. KV cache
  2. Prompt / prefix cache
  3. Distributed cache
Infrastructure

13.3 The golden rule

Don't ask:

"Where can I add caching?"

Ask:

"Where am I repeatedly doing expensive work that produces a reusable result?"

That shift in thinking is much more useful. Your AI system might repeatedly perform the same:

  • token processing
  • prompt processing
  • embedding
  • retrieval
  • reranking
  • tool call
  • API call
  • conversation reconstruction
  • agent step
  • answer

Each one is a potential caching opportunity.

14. The full picture

14.1 What caching buys you, and what it costs

A mature AI application can eventually run every layer from the architecture above. The result is not simply a faster application. A well-designed cache architecture gives you:

  1. AI system
  2. Lower latencyLower costLower load
  3. Better scale

But caching introduces its own engineering problems:

  • invalidation
  • freshness
  • consistency
  • security
  • multi-tenancy
  • versioning
  • memory
  • eviction
  • observability

That's why production AI caching is not about adding Redis somewhere in the architecture. It is about designing which work can safely be reused, for whom, for how long, and under which version of the system.

14.2 Part 1 takeaway

The 10 core caches are:

  1. KV cache
  2. Prompt cache
  3. Semantic cache
  4. Embedding cache
  5. Retrieval / vector search cache
  6. Reranking cache
  7. Tool result cache
  8. API response cache
  9. Conversation / context cache
  10. Agent state cache

And they operate at different layers:

Layer Caches
Model KV cache, prompt cache
RAG Embedding cache, retrieval cache, reranking cache
Agents Tool result cache, agent state cache
Conversation Context cache, semantic cache
Infrastructure API response cache

The most important production mindset is simple:

Cache the work, not just the answer.

Because in an AI system, the expensive work often happens before the final answer is generated.

Next: Part 2, Agentic AI caching covers tools, workflows, state, checkpoints, idempotency and MCP.