The AI Caching Playbook, Part 2: Agentic AI Caching
Caching for agents: tool results, workflow steps, state, checkpoints, MCP tools and resources, idempotency, and keeping it all safe.

On this page
A production agent shouldn't repeatedly perform work it has already completed.
This is Part 2 of the AI Caching Playbook. Part 1 covered the 10 core caches in the AI stack. This part goes deeper into agents: tools, workflows, state and MCP.
Agentic AI systems are fundamentally different from simple LLM applications.
A traditional LLM application often looks like:
- User
- Prompt
- LLM
- Response
A RAG application adds retrieval:
- User
- Query
- Retriever
- Documents
- LLM
- Response
An agent can be much more complicated:
- User
- Agent
- Plan
- Tool call
- Tool result
- Reason
- Another tool
- Another result
- Decision
- More tools
- Final answer
And this process may involve:
- APIs
- databases
- web searches
- MCP tools
- internal services
- other agents
- LLM calls
- long-running workflows
- retries
- checkpoints
- intermediate results
Without caching, an agent can repeatedly perform the same expensive work. For example, an agent handling a customer request might:
- Search customer
- Fetch orders
- Fetch product details
- Check inventory
- Calculate pricing
- Generate response
Now imagine the agent retries because of a temporary failure. It may do all of it again:
- Search customeragain
- Fetch ordersagain
- Fetch productagain
- Check inventoryagain
- Calculate pricingagain
That's unnecessary. A production-grade agent should be able to say:
"I already completed this step. I can reuse the result."
That's where agentic caching becomes important.
1. Tool result cache
1.1 What is it?
Tool calls are one of the biggest caching opportunities in agentic systems.
Suppose an agent has get_customer(), get_orders(), get_inventory(), get_weather(), search_web() and get_exchange_rate(). The agent might call get_customer(customer_id=123) several times during the same task. If the customer information hasn't changed, making the same API call again provides little value.
Instead:
- Agent
- Tool call
- Tool cache
- Hitreturn existing resultMissexecute tool
- Store result
1.2 Basic implementation
Using Redis:
import json
import redis.asyncio as redis
redis_client = redis.Redis(host="localhost", port=6379, decode_responses=True)
async def cached_tool_call(tool_name: str, arguments: dict, tool_fn, ttl: int = 300):
cache_key = f"tool:{tool_name}:{json.dumps(arguments, sort_keys=True)}"
cached = await redis_client.get(cache_key)
if cached:
return json.loads(cached)
result = await tool_fn(**arguments)
await redis_client.setex(cache_key, ttl, json.dumps(result))
return resultUsage:
customer = await cached_tool_call(
tool_name="get_customer",
arguments={"customer_id": "123"},
tool_fn=get_customer,
ttl=600,
)1.3 Production cache key
Don't key on the tool name alone. A tool can be called with different arguments.
Bad:
key = "get_customer"Better: combine the tool name + arguments + tool version + tenant ID. For example:
import hashlib
import json
def tool_cache_key(tenant_id: str, tool_name: str, arguments: dict, tool_version: str):
raw = json.dumps(
{
"tenant_id": tenant_id,
"tool": tool_name,
"arguments": arguments,
"version": tool_version,
},
sort_keys=True,
)
return hashlib.sha256(raw.encode()).hexdigest()2. Function call cache
2.1 What is it?
Tool caching and function-call caching are closely related, but function-call caching can be applied at a lower level of abstraction. Consider:
- LLM
- Function call
- get_exchange_rate()
- API
If the same function arguments produce the same result for a reasonable period, cache the function execution.
2.2 Example
import time
from functools import wraps
def cached(ttl=300):
cache = {}
def decorator(fn):
@wraps(fn)
def wrapper(*args, **kwargs):
key = (fn.__name__, args, tuple(sorted(kwargs.items())))
now = time.time()
if key in cache:
value, timestamp = cache[key]
if now - timestamp < ttl:
return value
result = fn(*args, **kwargs)
cache[key] = (result, now)
return result
return wrapper
return decoratorUsage:
@cached(ttl=60)
def get_exchange_rate(from_currency, to_currency):
return external_exchange_api(from_currency, to_currency)Now get_exchange_rate("USD", "INR") doesn't necessarily hit the external API every time.
2.3 Don't cache everything
This is critical. Consider create_payment(), delete_user(), place_order() and transfer_money().
These are side-effecting operations, and blindly caching them can be dangerous. place_order() should not be treated like a read such as get_customer().
For mutations, you usually need idempotency, which we'll cover later.
3. Agent state cache
3.1 What is it?
An agent isn't just producing an answer. It has state. For example:
{
"task_id": "task_123",
"goal": "Research company X",
"current_step": "financial_analysis",
"completed_steps": ["company_search", "financial_report"],
"results": {
"company": "...",
"revenue": "..."
}
}This state allows the agent to resume execution.
3.2 With and without state
- Step 1
- Step 2
- Step 3
- Crash
- Start again from step 1
- Step 1 ✓
- Step 2 ✓
- Step 3 ✓
- Crash
- Load state
- Step 4
3.3 Redis implementation
import json
async def save_agent_state(task_id: str, state: dict, ttl: int = 3600):
await redis_client.setex(f"agent:state:{task_id}", ttl, json.dumps(state))
async def load_agent_state(task_id: str):
data = await redis_client.get(f"agent:state:{task_id}")
if not data:
return None
return json.loads(data)3.4 Agent loop
state = await load_agent_state(task_id)
if not state:
state = {"completed_steps": [], "results": {}}
if "search" not in state["completed_steps"]:
result = await search_company()
state["results"]["company"] = result
state["completed_steps"].append("search")
await save_agent_state(task_id, state)The important part is:
Persist state after meaningful checkpoints, not only when the entire task finishes.
4. Workflow and step cache
4.1 What is it?
Agent workflows frequently contain deterministic steps. Consider a research agent:
- Search company
- Extract revenue
- Calculate growth
- Find competitors
- Generate report
If step 3 has already completed successfully, don't execute it again.
4.2 Step cache architecture
- Workflow
- Step
- Step cache
- Hitreturn resultMissexecute step
- Store in cache
4.3 Implementation
async def execute_step(workflow_id: str, step_name: str, input_data: dict, step_fn):
key = make_step_key(workflow_id, step_name, input_data)
cached = await redis_client.get(key)
if cached:
return json.loads(cached)
result = await step_fn(input_data)
await redis_client.setex(key, 3600, json.dumps(result))
return result4.4 Why this is powerful
Imagine the workflow fails at step 5 and has to be retried:
| Step | First run | Retry without step cache | Retry with step cache |
|---|---|---|---|
| 1 | 5 s | 5 s | cached |
| 2 | 8 s | 8 s | cached |
| 3 | 12 s | 12 s | cached |
| 4 | 20 s | 20 s | cached |
| 5 | 10 s | 10 s | 10 s |
| Total | 55 s | 55 s | ≈ 10 s |
This becomes extremely important for long-running agents.
5. MCP tool cache
5.1 What is it?
Model Context Protocol (MCP) systems introduce another caching opportunity. An agent might reach search, a database, files, a CRM and internal APIs through MCP servers. Repeated MCP calls can add unnecessary latency.
Conceptually:
- Agent
- MCP client
- MCP tool cache
- MCP server
- External system
5.2 Cache by tool and arguments
def mcp_cache_key(server, tool, arguments, version):
payload = {
"server": server,
"tool": tool,
"arguments": arguments,
"version": version,
}
raw = json.dumps(payload, sort_keys=True)
return hashlib.sha256(raw.encode()).hexdigest()Then:
async def call_mcp_tool(server, tool, arguments):
key = mcp_cache_key(server, tool, arguments, version="v1")
cached = await redis_client.get(f"mcp:{key}")
if cached:
return json.loads(cached)
result = await actual_mcp_call(server, tool, arguments)
await redis_client.setex(f"mcp:{key}", 300, json.dumps(result))
return result5.3 Be careful with MCP caching
Not every MCP resource should be cached equally:
| Data | Cache policy |
|---|---|
| Static documentation | Long TTL |
| Product metadata | Medium TTL |
| Live inventory | Short TTL |
| Financial transactions | Don't blindly cache |
| User permissions | Very careful |
The cache policy belongs to the data semantics, not just the tool name.
6. MCP resource cache
6.1 What is it?
MCP isn't only about tools. Resources can also represent data that agents access repeatedly, for example file://company/policy.pdf, db://customers/123 or docs://product/api.
A resource cache avoids fetching the same resource again and again.
6.2 Example
async def get_mcp_resource(resource_uri: str, fetch_fn, ttl: int = 600):
key = "mcp:resource:" + hashlib.sha256(resource_uri.encode()).hexdigest()
cached = await redis_client.get(key)
if cached:
return json.loads(cached)
resource = await fetch_fn(resource_uri)
await redis_client.setex(key, ttl, json.dumps(resource))
return resource6.3 Version-aware resources
If resources have versions, keying on the version too is safer:
- Any resource: resource URI + resource version
- Documents: document ID + document version
- APIs: resource ID +
updated_at
7. API response cache
7.1 What is it?
Agents frequently call external services: weather, stock prices, CRM, search, maps. External calls introduce:
- latency
- rate limits
- cost
- network failures
Caching can reduce all four.
7.2 Production implementation
import json
async def cached_api_response(key: str, fetch_fn, ttl: int = 300):
cache_key = f"api:{key}"
cached = await redis_client.get(cache_key)
if cached:
return json.loads(cached)
result = await fetch_fn()
await redis_client.setex(cache_key, ttl, json.dumps(result))
return resultUsage:
weather = await cached_api_response(
key="weather:bengaluru",
fetch_fn=lambda: fetch_weather("bengaluru"),
ttl=300,
)7.3 Stale-while-revalidate
For some applications, you don't need perfectly fresh data. You can serve a slightly stale result while refreshing the cache in the background:
- Request
- Cache
- FreshreturnStalereturn now, refresh in backgroundMissingfetch
This can dramatically reduce latency.
8. Agent memory cache
8.1 What is it?
Agent memory can exist at multiple levels:
| Short-term memory | Long-term memory |
|---|---|
| Current conversation | User preferences |
| Current task | Historical facts |
| Recent tool results | Learned context |
Not all memory needs to be fetched from a database every time.
8.2 Example
async def get_user_memory(user_id: str):
key = f"memory:user:{user_id}"
cached = await redis_client.get(key)
if cached:
return json.loads(cached)
memory = await database.get_user_memory(user_id)
await redis_client.setex(key, 1800, json.dumps(memory))
return memoryNow the agent can access frequently used memory quickly.
8.3 Memory invalidation
Suppose the user changes their preferred language from English to Tamil. Your memory cache must be invalidated:
await redis_client.delete(f"memory:user:{user_id}")Otherwise the agent may keep using stale information.
9. Planning and reasoning cache
9.1 What is it?
This is one of the more advanced forms of agent caching. Suppose an agent repeatedly receives:
"Analyze this invoice and extract: vendor, total, tax and due date."
The high-level plan may be reusable:
- Read invoice
- Extract fields
- Validate fields
- Return structured result
Instead of asking the LLM to generate the same plan every time:
- User request
- Planning LLM
- Plan
you can potentially cache stable plans.
9.2 Architecture
- Task
- Normalize intent
- Planning cache
- Hitreturn existing planMisscall planning LLM
- Store in cache
9.3 Example
async def get_plan(task):
normalized = normalize_task(task)
key = "plan:" + hashlib.sha256(normalized.encode()).hexdigest()
cached = await redis_client.get(key)
if cached:
return json.loads(cached)
plan = await planning_model(task)
await redis_client.setex(key, 3600, json.dumps(plan))
return plan9.4 Important warning
Reasoning plans are more sensitive than deterministic data. A cached plan may become invalid when:
- tools change
- permissions change
- available data changes
- system instructions change
- agent capabilities change
So include these versions in the cache key: task + agent version + toolset version + prompt version.
10. Task and job cache
10.1 What is it?
Long-running agents frequently run asynchronous jobs. For example:
- User
- "Analyze this 500-page document"
- Task ID
- Background agent
The user might refresh the page. You don't want to start another job. Instead:
- POST /analyze
- Task cache
- Existing taskreturn task_idNew taskstart agent
10.2 Idempotent task creation
async def create_task(user_id, request):
request_hash = hashlib.sha256(
json.dumps(request, sort_keys=True).encode()
).hexdigest()
key = f"task:{user_id}:{request_hash}"
existing = await redis_client.get(key)
if existing:
return json.loads(existing)
task = await start_agent_task(user_id, request)
await redis_client.setex(key, 3600, json.dumps(task))
return taskNow when a user clicks twice, the same request maps to the same task, instead of creating task 1 and task 2.
11. Idempotency cache
11.1 Caching is not idempotency
This deserves special attention. Consider POST /payments.
If the network times out, the client may retry, and the server receives payment request #1 and payment request #2. If both execute, the customer is charged ₹10,000 twice.
11.2 Idempotency key
The client generates a key, for example idempotency_key = abc123, and sends it with the request. The server:
async def process_payment(idempotency_key, payment_data):
cache_key = f"payment:{idempotency_key}"
existing = await redis_client.get(cache_key)
if existing:
return json.loads(existing)
result = await charge_payment(payment_data)
await redis_client.set(cache_key, json.dumps(result), ex=86400)
return resultNow retries become safe:
- Process payment
- Store result
- Look up key
- Return previous result
This is essential for agents that can trigger mutations.
12. Agent checkpoint cache
12.1 What is it?
Agent workflows can be extremely long. Consider:
- Search
- Download
- Parse
- Analyze
- Compare
- Validate
- Generate report
- Send email
If step 7 fails, restarting everything is wasteful. Checkpoint after important steps:
- Step 1 ✓checkpoint
- Step 2 ✓checkpoint
- Step 3 ✓checkpoint
- Step 4 ✓checkpoint
- Step 5 ✓checkpoint
- Step 6 ✓checkpoint
- Step 7 ✗
Then resume from the last checkpoint, straight into step 7.
12.2 Implementation
async def checkpoint(task_id, step, state):
key = f"checkpoint:{task_id}:{step}"
await redis_client.set(key, json.dumps(state))
async def load_checkpoint(task_id, step):
key = f"checkpoint:{task_id}:{step}"
data = await redis_client.get(key)
if not data:
return None
return json.loads(data)13. Sub-agent result cache
13.1 What is it?
Modern agents can delegate tasks:
- Main agent
- Research agentCoding agentData agentValidation agent
Suppose the research agent has already answered:
"Find the company's revenue for 2025."
The main agent shouldn't necessarily trigger another research agent for the exact same task.
13.2 Architecture
- Main agent
- Sub-agent request
- Sub-agent cache
- Hitreturn resultMissrun sub-agent
- Store in cache
13.3 Key design
Key on tenant + sub-agent type + task + input + agent version. For example:
def subagent_key(tenant_id, agent_name, task, input_data, version):
raw = json.dumps(
{
"tenant": tenant_id,
"agent": agent_name,
"task": task,
"input": input_data,
"version": version,
},
sort_keys=True,
)
return hashlib.sha256(raw.encode()).hexdigest()14. Multi-agent shared cache
14.1 What is it?
Now consider several agents working together:
- Main agent
- Research agentData agent
- Shared cache
Agents may need to share results. For example:
- Research agent
- Company revenue = $4.2B
- Shared cache
- Financial agent
The financial agent doesn't need to fetch the same information again.
14.2 But shared state introduces problems
You need to define:
- Who can write?
- Who can read?
- Who owns the data?
- How long is it valid?
- What version produced it?
- Can another tenant access it?
For enterprise systems, this becomes an authorization problem. A safe key might include tenant ID + namespace + resource ID + version.
15. Long-running agent cache
15.1 What is it?
Some agents run for seconds, some for minutes and some for hours: deep research, data analysis, document processing, code generation, report generation, monitoring.
You don't want the agent's entire state to live only in memory. Instead, the worker loops:
- Execute
- Checkpoint
- Save state
- Continue
If the worker dies, another one picks up where it stopped:
- Worker 1
- Crash
- Worker 2
- Load checkpoint
- Resume
16. Designing a production agent cache
16.1 Putting the pieces together
At this point, we can combine the pieces:
- User
- API gateway
- Task / idempotency cache
- Agent state
- Agent
- PlanningToolsSub-agent
- Plan cacheTool cacheSub-agent cache
- MCPvia toolsAPIvia tools
- MCP cacheAPI cache
- Resources
- Resource cache
16.2 Cache key design for agents
Cache key design becomes extremely important once multiple agents and tools are involved.
- A weak key:
"get_customer" - A better key: tenant ID + tool name + arguments
- A production-oriented key: tenant ID + user scope + tool name + arguments + tool version + agent version + data version
For example:
def make_agent_cache_key(
tenant_id,
user_scope,
component,
input_data,
component_version,
data_version=None,
):
payload = {
"tenant_id": tenant_id,
"user_scope": user_scope,
"component": component,
"input": input_data,
"component_version": component_version,
"data_version": data_version,
}
raw = json.dumps(payload, sort_keys=True)
return hashlib.sha256(raw.encode()).hexdigest()16.3 TTL strategy for agentic systems
Different agent components need different TTLs:
| Component | Example TTL |
|---|---|
| Weather tool | 5 min |
| Customer profile | 10 min |
| Search results | 5–30 min |
| Product metadata | 1 hour |
| Agent plan | 1 hour |
| Static documentation | 24 hours |
| Tool result | Tool-dependent |
| MCP resource | Resource-dependent |
| Sub-agent result | Task-dependent |
| Agent state | Until task completion |
| Checkpoint | Until task completion |
| Payment result | Idempotency window |
These are examples, not universal values. The correct TTL depends on freshness requirements, data volatility, business impact, cost and failure-recovery requirements.
16.4 Cache invalidation in agentic systems
This is where agent caching gets complicated. Suppose the agent reads a customer through a tool backed by a database:
- Agent
- Tool
- Database
In the database, the customer's status changes from ACTIVE to SUSPENDED. But the agent cache still says ACTIVE. Now the agent acts on stale information.
TTL-based
- Cache
- 10 minutes
- Expire
Simple, but not always precise.
Event-based
- Database update
- Event
- Invalidate cache
More accurate.
Version-based
Keep a version on the record, for example customer_version = 42, and put it in the cache key: customer:123:v42. When the customer changes, customer_version becomes 43, and the old cache entry naturally becomes irrelevant.
Explicit invalidation
await redis_client.delete(f"customer:{customer_id}")Useful for important state changes.
16.5 Caching and agent safety
Caching introduces a new attack surface. Consider:
- User A"What is customer 123's balance?"
- Response cached
- User B"Show me customer 123's balance"
- Cache hit
- Private data leak
If the cache isn't authorization-aware, User B sees data they may not be allowed to see. So the cache key should combine tenant + user permissions + resource + query + version.
16.6 Never cache sensitive results blindly
Be particularly careful with:
- passwords
- authentication and access tokens
- payment information
- private customer information
- personal data
- healthcare information
- financial information
- secrets
Caching can turn a normal application bug into a cross-user data leak.
16.7 Cache stampede
Another important production problem is the cache stampede. Imagine a popular cache entry expires:
- Cache expires
- 10,000 requests
- 10,000 API calls
You just destroyed the purpose of caching.
The fix is a lock, so only one worker refreshes the entry:
import asyncio
import json
async def get_with_lock(key, fetch_fn, ttl=300):
cached = await redis_client.get(key)
if cached:
return json.loads(cached)
lock_key = f"{key}:lock"
acquired = await redis_client.set(lock_key, "1", nx=True, ex=30)
if acquired:
try:
result = await fetch_fn()
await redis_client.setex(key, ttl, json.dumps(result))
return result
finally:
await redis_client.delete(lock_key)
# Another worker is refreshing. Wait briefly and retry.
await asyncio.sleep(0.1)
return await get_with_lock(key, fetch_fn, ttl)This prevents thousands of workers from performing the same expensive operation.
16.8 Cache observability
Production agents need cache telemetry. At minimum, track cache hits, misses, errors, evictions, latency, size, stale reads and refresh count.
For agents, also track tool calls avoided, agent steps avoided, LLM calls avoided, API calls avoided, tokens saved, agent runtime saved and cost saved.
An agent cache dashboard might show:
| Metric | Value |
|---|---|
| Tool cache hit rate | 72.4% |
| MCP cache hit rate | 61.8% |
| Agent state hit rate | 98.1% |
| Plan cache hit rate | 84.3% |
| API calls avoided | 42K |
| Agent steps avoided | 18K |
| LLM calls avoided | 9K |
| Tokens saved | 14.2M |
| Estimated cost saved | $980 |
17. Cache, state and idempotency
17.1 Caching vs idempotency
These concepts are related but different.
- Caching: "I already have the answer."
- Idempotency: "I already processed this operation."
GET customer is a good fit for caching. POST payment requires idempotency.
This distinction is extremely important when building agents that can take actions.
17.2 Caching vs state
Another important distinction:
- Cache: temporary, reusable data, such as tool results, search results and API responses.
- State: information required to continue execution, such as the current task, completed steps, the agent plan, checkpoints and workflow status.
State may use the same infrastructure as a cache, such as Redis, but conceptually they are different. For critical state, you may want durable storage in addition to Redis:
- Agent
- Redisfast stateDatabasedurable state
17.3 A shared cache layer
Instead of letting every component (tools, MCP, the agent, API clients) talk to Redis directly and implement caching differently, create a common abstraction:
import json
class Cache:
def __init__(self, redis_client):
self.redis = redis_client
async def get(self, key):
value = await self.redis.get(key)
if value is None:
return None
return json.loads(value)
async def set(self, key, value, ttl):
await self.redis.setex(key, ttl, json.dumps(value))
async def delete(self, key):
await self.redis.delete(key)Now cache = Cache(redis_client) can be shared across the agent, tool layer, MCP layer, workflow layer, API layer and memory layer.
17.4 Recommended agent cache architecture
For a serious production system:
- Agent
- PlanningToolsWorkflow
- Plan cacheTool cacheStep cache
- MCP cachevia toolsAPI cachevia tools
- Agent state
- Redisfast storePostgreSQLdurable store
This gives you fast access, recovery, deduplication, lower cost, lower latency and better scalability.
18. Before you ship
18.1 Production checklist
Before deploying agent caching, ask:
Tool caching
- Is this tool safe to cache?
- What is the TTL?
- Are arguments included in the key?
- Is the tool version included?
- Is tenant isolation enforced?
Agent state
- Can the agent resume after failure?
- Are checkpoints persisted?
- Is state durable?
- Can two workers modify the same state?
MCP
- Which tools are cacheable?
- Which resources are cacheable?
- Are resource versions tracked?
- Are permissions included?
APIs
- Is the API response deterministic?
- Does the provider allow caching?
- What is the freshness requirement?
- What happens when the cache is stale?
Idempotency
- Are mutations protected?
- Are retries safe?
- Is there an idempotency key?
- Can duplicate agent actions happen?
Operations
- Cache hit and miss metrics?
- Cache latency?
- Cache size?
- Eviction policy?
- Stampede protection?
- Invalidation strategy?
- Security audit?
19. The full picture
19.1 The core principle
Agentic caching isn't about putting Redis between your agent and everything else. It's about identifying repeated work. An agent may repeatedly perform the same:
- Plan
- Tool call
- API request
- MCP resource fetch
- Workflow step
- Sub-agent task
A production system should recognize "I have already done this" and safely reuse the result.
19.2 Part 2 takeaway
The major caching layers for agentic AI are:
- Tool result cache
- Function call cache
- Agent state cache
- Workflow / step cache
- MCP tool cache
- MCP resource cache
- API response cache
- Agent memory cache
- Planning / reasoning cache
- Task / job cache
- Idempotency cache
- Agent checkpoint cache
- Sub-agent result cache
- Multi-agent shared cache
- Long-running agent cache
The most important production distinction is:
| Concept | What it does |
|---|---|
| Cache | Reuse previous work |
| State | Resume previous work |
| Idempotency | Prevent duplicate work |
And the overall architecture becomes:
- User
- Agent
- PlanningToolsWorkflow
- Plan cacheTool cacheStep cache
- MCPvia toolsAPIsvia tools
- MCP cacheAPI cache
- Resources
- Agent state
- RedisDurable DB
The best agent isn't the one that calls the most tools. It's the one that knows when it doesn't need to call them again.
Next: Part 3, Conversational AI caching covers context, sessions, memory and responses.