HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 3: Conversational AI Caching

Caching for chat: conversation history, sessions, summaries, user memory, context fingerprints and responses, without leaking or going stale.

Haribaskar Dhanabalan17 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

A good conversational agent remembers what matters without repeatedly paying to process everything.

In Part 1, we looked at the core caching layers used across AI systems. In Part 2, we moved into agentic AI caching: tool results, workflow steps, agent state, MCP resources, checkpoints and idempotency.

Now we move to another important AI architecture: conversational AI.

A conversational agent has a different problem. A normal request might look like:

  1. User
  2. Prompt
  3. LLM
  4. Response

But a real conversational application looks more like:

  1. User
  2. Conversation
  3. Session
  4. Recent messages
  5. Conversation summary
  6. User memory
  7. System instructions
  8. Tools
  9. LLM
  10. Response

As the conversation grows from 10 messages to 50, 200 and 1,000, the system faces several problems:

  • more tokens
  • higher latency
  • higher inference cost
  • larger context windows
  • more database reads
  • more repeated context
  • more difficult context management
  • more opportunities for stale information
  • more complicated memory management

A production conversational system therefore needs more than a simple message database. It needs a context and memory caching strategy.

1. The conversational caching stack

A production conversational AI system can have several cache layers:

  1. User
  2. Session cache
  3. Recent chat cacheUser memory cacheContext cache
  4. Prompt cache
  5. LLM
  6. Response / semantic cache

The goal isn't simply to cache the conversation. The goal is:

Build the minimum useful context required for the next response, as efficiently as possible.

2. Conversation history cache

2.1 What is it?

The simplest and most important layer is the conversation history cache. Consider a conversation:

User: My order number is ORD-123.

Assistant: Got it.

User: What is its status?

Assistant: Your order has shipped.

User: When will it arrive?

The agent needs the previous context. A naive implementation queries the database every time:

  1. Request
  2. Database
  3. Fetch entire conversation
  4. Build prompt
  5. LLM

For high-traffic applications, this can become expensive. Instead:

  1. Request
  2. Conversation cache
  3. Hitrecent messagesMissload from database
  4. Store in cache

2.2 Redis implementation

import json
 
import redis.asyncio as redis
 
redis_client = redis.Redis(host="localhost", port=6379, decode_responses=True)
 
 
async def save_messages(conversation_id: str, messages: list, ttl: int = 3600):
    key = f"conversation:{conversation_id}"
    await redis_client.setex(key, ttl, json.dumps(messages))
 
 
async def get_messages(conversation_id: str):
    key = f"conversation:{conversation_id}"
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    messages = await load_from_database(conversation_id)
    await save_messages(conversation_id, messages)
    return messages

2.3 Don't cache unlimited history

A common mistake is to cache the entire conversation and then send the entire conversation to the LLM.

Suppose the conversation contains 50,000 tokens. The fact that you retrieved it quickly doesn't mean the LLM can process it cheaply. You need a second layer:

  1. Stored conversation
  2. Context builder
  3. Relevant context
  4. LLM

This distinction is important:

Caching conversation history and optimizing model context are two different problems.

3. Session cache

3.1 What is it?

Conversation history and session state are related, but they aren't the same. A session might contain:

{
  "session_id": "sess_123",
  "user_id": "user_456",
  "conversation_id": "conv_789",
  "language": "en",
  "current_page": "/checkout",
  "active_task": "order_tracking",
  "authenticated": true
}

This information may be needed on almost every request. Instead of loading it from the database every time, read it from a session cache:

  1. Request
  2. Session cache
  3. Session

3.2 Implementation

async def save_session(session_id: str, session: dict, ttl: int = 1800):
    await redis_client.setex(f"session:{session_id}", ttl, json.dumps(session))
 
 
async def get_session(session_id: str):
    data = await redis_client.get(f"session:{session_id}")
 
    if not data:
        return None
 
    return json.loads(data)

3.3 Session TTL

Sessions are usually short-lived compared with long-term memory, for example 30 minutes before they expire. But this depends on your application:

  • Customer support agent: 15–60 minutes might be reasonable.
  • Long-running enterprise workflow: hours or days may be required.

4. Context cache

4.1 What is it?

This is where conversational AI becomes more interesting.

A conversation can contain thousands of messages, but the model doesn't necessarily need all of them. Instead, you can build a context package from:

  • recent messages
  • the summary
  • important facts
  • user preferences
  • the current task
  • relevant historical messages
  1. Conversation
  2. Context package
  3. LLM

You can cache this package.

4.2 Example

async def build_context(conversation_id: str):
    messages = await get_messages(conversation_id)
    summary = await get_summary(conversation_id)
    memory = await get_user_memory(conversation_id)
 
    return {
        "summary": summary,
        "memory": memory,
        "recent_messages": messages[-10:],
    }

Then cache it under a versioned key:

context_key = f"context:{conversation_id}:{context_version}"

4.3 Why version context?

Suppose context V1 was built from the user's preferences and the conversation summary. Then the user says:

"Actually, my preferred language is Tamil."

The context changes. If the cache key stays context:123, you could keep using stale context. Instead, move to context:123:v2, or invalidate the entry explicitly.

5. Conversation summary cache

5.1 What is it?

Long conversations need summarization. Turning 100 messages into a conversation summary costs an LLM call, and generating that summary again and again is wasteful. Cache it:

  1. Messages 1–100
  2. Summary
  3. Cache
Messages 1–100 are summarized once and cached

New messages, say 101–110, are then sent as fresh context alongside the cached summary.

5.2 Basic implementation

async def get_conversation_summary(conversation_id: str):
    key = f"summary:{conversation_id}"
 
    cached = await redis_client.get(key)
    if cached:
        return cached
 
    messages = await get_messages(conversation_id)
    summary = await summarize_with_llm(messages)
 
    await redis_client.setex(key, 3600, summary)
    return summary

5.3 Better: version the summary

A summary should represent a specific section of the conversation, for example summary_version = 12 covering messages 1–150.

When messages 151–170 arrive, the old summary of messages 1–150 is still useful. You can build summaries in segments:

Summary Messages covered
Summary 1 1–150
Summary 2 151–300

This is better than repeatedly summarizing the entire conversation.

6. Context summarization cache

6.1 What is it?

A production conversational system often needs hierarchical context:

  1. Raw messages
  2. Message-level history
  3. Conversation summary
  4. Long-term user memory
  5. Current task context

Instead of sending 10,000 messages, the model might receive:

  1. System instructions
  2. Long-term memory
  3. Conversation summary
  4. Last 10 messages
  5. Relevant historical messages

This dramatically reduces context size.

6.2 Example context builder

async def build_llm_context(conversation_id, user_id):
    summary = await get_summary(conversation_id)
    memory = await get_user_memory(user_id)
    recent_messages = await get_recent_messages(conversation_id, limit=10)
 
    return [
        {"type": "summary", "content": summary},
        {"type": "memory", "content": memory},
        {"type": "messages", "content": recent_messages},
    ]

7. User memory cache

7.1 What is it?

Conversational agents often keep long-term information about a user, such as:

  • prefers concise answers
  • works with Python
  • is building a RAG system
  • prefers technical examples

This memory may live in PostgreSQL, Redis, a vector database, a document store or a dedicated memory service. If the same user interacts frequently, fetching their memory from the database every time is unnecessary.

7.2 Cache architecture

  1. User
  2. Memory request
  3. Memory cache
  4. Hitreturn memoryMissload from database
  5. Store in cache

7.3 Implementation

async def get_user_memory(user_id: str):
    key = f"user-memory:{user_id}"
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    memory = await database.load_memory(user_id)
 
    await redis_client.setex(key, 1800, json.dumps(memory))
    return memory

7.4 Memory invalidation

Suppose the user says:

"I don't work with Python anymore. I use Go now."

Your application needs to update memory. After updating the durable store, delete the cached copy:

await database.update_memory(user_id, new_memory)
 
await redis_client.delete(f"user-memory:{user_id}")

Otherwise the agent may keep assuming Python instead of Go.

8. Semantic response cache

8.1 What is it?

This is where caching becomes extremely powerful. Suppose users ask:

  1. "What are your refund rules?"
  2. "Can I get my money back?"
  3. "How does the refund process work?"

These are different strings, but potentially the same intent. A normal cache matches the exact query string, so it will miss. A semantic cache can reuse the previous response:

  1. Query
  2. Embedding
  3. Similarity search
  4. Cached response

8.2 Architecture

  1. User query
  2. Embedding
  3. Semantic cache
  4. Hitreturn responseMisscall the LLM
  5. Store in cache

8.3 Simplified implementation

import numpy as np
from sentence_transformers import SentenceTransformer
 
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")
 
semantic_cache = []
 
 
def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
 
 
def semantic_lookup(query, threshold=0.90):
    query_embedding = embedding_model.encode(query)
 
    for item in semantic_cache:
        score = cosine_similarity(query_embedding, item["embedding"])
        if score >= threshold:
            return {"response": item["response"], "score": float(score)}
 
    return None

Store:

def store_semantic_response(query, response):
    embedding = embedding_model.encode(query)
 
    semantic_cache.append({
        "query": query,
        "embedding": embedding,
        "response": response,
    })

8.4 Production semantic cache

A Python list isn't suitable for production. You would typically use Redis with vector search, a vector database, or a dedicated semantic caching layer.

The cache record might look like:

{
  "tenant_id": "tenant_123",
  "query": "What is your refund policy?",
  "embedding": "...",
  "response": "...",
  "model_version": "model-v3",
  "prompt_version": "prompt-v8",
  "knowledge_version": "kb-v17",
  "created_at": "...",
  "expires_at": "..."
}

9. System prompt and prefix cache

9.1 What is it?

Conversational agents often have large system instructions:

You are an enterprise support agent.
 
Rules:
...
Security policies:
...
Company policies:
...
Tool definitions:
...
Response formatting:
...

This may be thousands of tokens. If the system prompt stays the same across requests, prompt or prefix caching can avoid processing the same prefix again, where the model-serving infrastructure or provider supports it.

9.2 Cache-friendly structure

Prefer a static prefix followed by a dynamic suffix, rather than constantly changing the beginning of the prompt:

  1. System instructions
  2. Company policies
  3. Tool definitions
  4. Stable context
Static prefix: the same on every request
  1. Conversation
  2. User query
Dynamic suffix: changes on every request

9.3 Why prompt ordering matters

Imagine request 1 is SYSTEM, POLICY, TOOLS, USER QUERY and request 2 is SYSTEM, POLICY, TOOLS, DIFFERENT USER QUERY. The common prefix is large, which is exactly what prefix caching can reuse.

But if you build the prompt as USER QUERY, SYSTEM, POLICY, TOOLS and the query changes every time, the beginning of the prompt changes on every request. You lose the opportunity for prefix reuse.

10. Conversation state cache

10.1 What is it?

Conversation state is more than messages. Consider an e-commerce agent:

{
  "intent": "product_search",
  "selected_product": "iphone",
  "selected_color": "black",
  "selected_storage": "256GB",
  "cart_items": 2,
  "checkout_stage": "payment"
}

The agent needs this state repeatedly. Instead of rebuilding it from the conversation history every time (conversation → LLM → infer state), store the state explicitly.

10.2 Example

async def save_conversation_state(conversation_id, state):
    await redis_client.setex(f"state:{conversation_id}", 3600, json.dumps(state))
 
 
async def get_conversation_state(conversation_id):
    data = await redis_client.get(f"state:{conversation_id}")
 
    if not data:
        return {}
 
    return json.loads(data)

10.3 Explicit state beats reconstructing everything

Instead of asking the model on every turn:

"Based on the entire conversation, what product did the user select?"

maintain:

{
  "selected_product": "MacBook Pro",
  "selected_storage": "1TB"
}

This gives you fewer tokens, lower latency, lower cost and more deterministic behaviour.

11. Personalization cache

11.1 What is it?

Personalized conversational systems often fetch the same user data again and again: preferences, language, timezone, subscription, product preferences and account information. These can be cached separately.

11.2 Example

async def get_profile(user_id):
    key = f"profile:{user_id}"
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    profile = await database.get_profile(user_id)
 
    await redis_client.setex(key, 1800, json.dumps(profile))
    return profile

Then every conversation request draws on several small caches:

  1. Conversation request
  2. Profile cacheMemory cacheSession cacheConversation cache
  3. Context builder
  4. LLM

12. Response cache

12.1 What is it?

Sometimes the exact same response can be reused. "What are your business hours?" might be asked thousands of times. If the answer is stable (9 AM – 6 PM, Monday to Friday), there may be no reason to invoke an LLM every time.

12.2 Exact response cache

import hashlib
 
 
async def get_response(query: str):
    key = "response:" + hashlib.sha256(query.encode()).hexdigest()
 
    cached = await redis_client.get(key)
    if cached:
        return cached
 
    response = await call_llm(query)
 
    await redis_client.setex(key, 3600, response)
    return response

12.3 But don't blindly cache LLM responses

LLM responses depend on:

  • the model
  • the prompt
  • system instructions
  • context
  • tools
  • knowledge
  • the user
  • temperature and other generation settings

So hash(query) is rarely enough for a serious conversational system. A better conceptual key is user or tenant scope + query + model version + prompt version + knowledge version + tool configuration + generation configuration.

13. Context-aware response cache

13.1 What is it?

In conversational applications, even the same query can have different answers. Compare two conversations:

Conversation A Conversation B
User: I'm asking about order 123. User: I'm asking about order 456.
User: Where is it? User: Where is it?

The query "Where is it?" is identical. The context isn't. A query-only cache would be dangerous.

13.2 Key on the context too

Instead, the cache key can include a context fingerprint:

def response_cache_key(user_id, query, context_version, model_version, prompt_version):
    payload = {
        "user_id": user_id,
        "query": query,
        "context_version": context_version,
        "model_version": model_version,
        "prompt_version": prompt_version,
    }
    raw = json.dumps(payload, sort_keys=True)
    return hashlib.sha256(raw.encode()).hexdigest()

14. Conversation context fingerprints

14.1 What is it?

A useful technique is to keep a version, or fingerprint, for the current conversation context. For example, a conversation might be at:

  • summary V4
  • memory V7
  • state V12
  • recent messages V30

Combine them into one fingerprint:

context_fingerprint = hash(summary_version + memory_version + state_version + recent_message_version)

Now query + context fingerprint + model version can become a response cache key. When the context changes, the old response naturally stops matching:

  1. Context V30
  2. New message
  3. Context V31

15. Multi-tenant conversation cache

15.1 Isolation first

This is one of the most important production considerations. Imagine:

  1. Tenant Aconversation cached
  2. Tenant Bsends a similar request
  3. Cache hit
  4. Tenant A's data returned to Tenant B

If your cache isn't properly isolated, Tenant B might receive Tenant A's data.

Never use:

key = f"conversation:{conversation_id}"

if IDs aren't globally unique, or if authorization isn't separately enforced. Prefer scoped keys:

key = (
    f"tenant:{tenant_id}:"
    f"user:{user_id}:"
    f"conversation:{conversation_id}"
)

For highly sensitive systems, also validate authorization before returning a cache hit.

16. Cache stampede in conversational systems

16.1 What is it?

Imagine a popular assistant has a cached company-policy response. At 10:00 AM the cache entry expires. At 10:00:01, 5,000 users ask the same question. Without protection:

  1. 5,000 requests
  2. 5,000 LLM calls

That's a cache stampede.

16.2 Locking

Let one worker regenerate the entry while the others wait briefly:

import asyncio
 
 
async def get_or_generate(key, generate_fn, ttl=3600):
    cached = await redis_client.get(key)
    if cached:
        return cached
 
    lock_key = f"{key}:lock"
    acquired = await redis_client.set(lock_key, "1", nx=True, ex=30)
 
    if acquired:
        try:
            result = await generate_fn()
            await redis_client.setex(key, ttl, result)
            return result
        finally:
            await redis_client.delete(lock_key)
 
    await asyncio.sleep(0.1)
    return await get_or_generate(key, generate_fn, ttl)

17. Stale-while-revalidate for conversations

17.1 What is it?

Some conversational data doesn't need to be perfectly fresh: product descriptions, the company FAQ, documentation, business hours, general policies. For these you can serve from the cache and refresh behind the scenes:

  1. Cached answer
  2. Freshreturn immediatelyStalereturn immediately, refresh in background

This gives you low latency with reasonable freshness.

18. Invalidation and dependencies

18.1 Cache invalidation

The hardest part of conversational caching is often invalidation. Consider user memory: the user's preferred language is English, and they change it to Tamil. You now need to invalidate the profile cache, the memory cache, the context cache and potentially the response cache.

That's a dependency chain:

  1. Profile
  2. Memory
  3. Context
  4. Response

Changing one layer can invalidate everything downstream.

18.2 Version-based invalidation

One powerful strategy is versioning. For example, profile_version = 12, memory_version = 7 and context_version = 18, with a response key of:

user + query + profile version + memory version + context version + model version

When memory changes, memory_version becomes 8 and the old response keys no longer match. This avoids explicitly deleting every dependent cache entry.

18.3 Cache dependencies

A useful production model is to think of caches as a dependency graph:

  1. User profile
  2. User memory
  3. Conversation context
  4. LLM response

If the user profile changes, the memory may become stale, so the context may become stale, so the response may become stale. Therefore:

Cache invalidation should follow data dependencies, not just individual cache entries.

19. Architecture

19.1 Conversation cache architecture

Putting everything together:

  1. User
  2. Session cache
  3. Conversation cacheUser memory cacheUser profile cache
  4. Context builder
  5. Context cache
  6. Prompt / prefix cache
  7. LLM
  8. Response cache

19.2 A production context builder

A practical conversational system should have a dedicated context builder:

async def build_context(tenant_id, user_id, conversation_id):
    session = await get_session(conversation_id)
    profile = await get_profile(user_id)
    memory = await get_user_memory(user_id)
    summary = await get_conversation_summary(conversation_id)
    recent_messages = await get_recent_messages(conversation_id, limit=10)
    state = await get_conversation_state(conversation_id)
 
    return {
        "session": session,
        "profile": profile,
        "memory": memory,
        "summary": summary,
        "recent_messages": recent_messages,
        "state": state,
    }

Then:

context = await build_context(tenant_id, user_id, conversation_id)
 
response = await llm.generate(context=context, query=user_query)

The key architectural principle is:

Don't let the LLM decide what historical data to load on every request. Build a deterministic context layer around it.

19.3 Context window management

Caching doesn't remove the context-window problem. Suppose a conversation is 100,000 tokens. Even if Redis retrieves it in 2 ms, sending it to the model can still be expensive.

So the context pipeline should look like:

  1. Full conversation
  2. Message store
  3. SummaryImportant factsUser memoryCurrent stateRecent messages
  4. Context selection
  5. LLM input

This is context optimization, not simply caching.

19.4 A complete request flow

A production conversational request, for "Can you check my order?", might look like:

  1. User"Can you check my order?"
  2. Session cache
  3. Conversation cache
  4. User memory cache
  5. Context builder
  6. Semantic cache
  7. Hitreturn responseMisshand to the agent
  8. Agent
  9. Tool → tool cache
  10. LLM
  11. Response cache

20. Deciding what to cache

20.1 What should be cached?

A practical decision table:

Data Cache? Typical strategy
Recent messages Yes Session TTL
Conversation summary Yes Versioned
User profile Yes Short / medium TTL
User preferences Yes Event invalidation
Long-term memory Yes Versioned
Conversation state Yes Session / task TTL
Static system prompt Yes Prefix caching
FAQ responses Yes Semantic cache
Dynamic account balance Carefully Very short TTL
Payment status Carefully Freshness-aware
Authentication tokens Usually no Secure token storage
Private sensitive data Carefully Strict isolation

20.2 What should not be cached blindly?

Be extremely careful with:

  • passwords
  • access tokens
  • payment information
  • private account information
  • healthcare information
  • financial information
  • authorization decisions
  • security-sensitive state

The question isn't:

"Can I cache it?"

The question is:

"What happens if this cached value is stale, or exposed to the wrong user?"

20.3 The most important cache key

For conversational AI, this is dangerous:

key = hash(query)

because "Where is it?" can mean different things in different conversations. A better key includes the context. Conceptually:

tenant + user + conversation context + query + model version + prompt version + knowledge version

For example:

import hashlib
import json
 
 
def make_response_cache_key(
    tenant_id,
    user_id,
    query,
    context_version,
    model_version,
    prompt_version,
    knowledge_version,
):
    payload = {
        "tenant": tenant_id,
        "user": user_id,
        "query": query,
        "context": context_version,
        "model": model_version,
        "prompt": prompt_version,
        "knowledge": knowledge_version,
    }
    raw = json.dumps(payload, sort_keys=True)
    return hashlib.sha256(raw.encode("utf-8")).hexdigest()

21. Observability

21.1 What to measure

A conversational cache should expose metrics. At minimum, the hit rate of each layer:

  • conversation_cache_hit_rate
  • session_cache_hit_rate
  • memory_cache_hit_rate
  • context_cache_hit_rate
  • semantic_cache_hit_rate
  • response_cache_hit_rate

Also track cache latency, cache size, evictions, stale reads, invalidations, context tokens saved, LLM calls avoided, tokens saved and cost saved.

A conversational AI cache dashboard might show:

Metric Value
Session hit rate 97.8%
Conversation hit rate 94.2%
Memory hit rate 91.5%
Context hit rate 76.4%
Semantic hit rate 63.8%
Context tokens saved 18.2M
LLM calls avoided 8.7K
Estimated cost saved $1,120
Stale reads 0.02%

22. Before you ship

22.1 Production checklist

Before deploying conversational caching, ask:

Conversation

  • Are recent messages cached?
  • Is conversation history stored durably?
  • Is the history window defined?
  • Is a summary strategy implemented?

Session

  • Is the session TTL defined?
  • Is session invalidation implemented?
  • Are sessions isolated by user and tenant?

Memory

  • Is long-term memory separate from the conversation?
  • Is memory versioned?
  • Is memory invalidation implemented?

Context

  • Does a context builder exist?
  • Is context size limited?
  • Are context versions tracked?
  • Is stale context detectable?

Responses

  • Is the semantic cache threshold tested?
  • Is the response cache context-aware?
  • Is the model version included?
  • Is the prompt version included?
  • Is the knowledge version included?

Security

  • Is tenant isolation enforced?
  • Is authorization checked before cache hits?
  • Is there a sensitive-data policy?
  • Is PII handled?
  • Are cache encryption and access controls in place?

Operations

  • Hit and miss metrics?
  • Latency metrics?
  • Token savings?
  • Cost savings?
  • Stale-read monitoring?
  • Cache stampede protection?

23. The full picture

23.1 The core principle

A conversational AI system should not treat its entire conversation as one giant prompt. Instead:

  1. Conversation
  2. Analyze
  3. Recent messagesSummaryUser memoryUser profileCurrent stateRelevant history
  4. Context package
  5. LLM

The goal is:

Remember more while sending less.

That's the fundamental idea behind conversational caching.

23.2 Part 3 takeaway

The important conversational caching layers are:

  1. Conversation history cache
  2. Session cache
  3. Context cache
  4. Conversation summary cache
  5. Context summarization cache
  6. User memory cache
  7. Semantic response cache
  8. System prompt / prefix cache
  9. Conversation state cache
  10. Personalization cache
  11. Response cache
  12. Context-aware response cache
  13. Context fingerprint cache
  14. Multi-tenant conversation cache
  15. Stale-while-revalidate cache

The most important distinction is:

Concept What it holds
Conversation history Everything that happened
Context Everything relevant right now
Memory What should be remembered across conversations
State What the agent currently knows or is doing
Cache What expensive work can safely be reused

And the production architecture becomes:

  1. Conversational agent
  2. Session cacheConversation cacheMemory cache
  3. Context builder
  4. Context cache
  5. Prompt / prefix cache
  6. LLM
  7. Response cache
  8. Semantic cache

The best conversational agent isn't the one that remembers everything. It's the one that knows what to remember, what to retrieve, what to cache, and what to leave out.

Next: Part 4, RAG caching covers queries, embeddings, retrieval, reranking, context and responses.