HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 2: Agentic AI Caching

Caching for agents: tool results, workflow steps, state, checkpoints, MCP tools and resources, idempotency, and keeping it all safe.

Haribaskar Dhanabalan16 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

A production agent shouldn't repeatedly perform work it has already completed.

This is Part 2 of the AI Caching Playbook. Part 1 covered the 10 core caches in the AI stack. This part goes deeper into agents: tools, workflows, state and MCP.

Agentic AI systems are fundamentally different from simple LLM applications.

A traditional LLM application often looks like:

  1. User
  2. Prompt
  3. LLM
  4. Response

A RAG application adds retrieval:

  1. User
  2. Query
  3. Retriever
  4. Documents
  5. LLM
  6. Response

An agent can be much more complicated:

  1. User
  2. Agent
  3. Plan
  4. Tool call
  5. Tool result
  6. Reason
  7. Another tool
  8. Another result
  9. Decision
  10. More tools
  11. Final answer

And this process may involve:

  • APIs
  • databases
  • web searches
  • MCP tools
  • internal services
  • other agents
  • LLM calls
  • long-running workflows
  • retries
  • checkpoints
  • intermediate results

Without caching, an agent can repeatedly perform the same expensive work. For example, an agent handling a customer request might:

  1. Search customer
  2. Fetch orders
  3. Fetch product details
  4. Check inventory
  5. Calculate pricing
  6. Generate response

Now imagine the agent retries because of a temporary failure. It may do all of it again:

  1. Search customeragain
  2. Fetch ordersagain
  3. Fetch productagain
  4. Check inventoryagain
  5. Calculate pricingagain
A retry without caching repeats every step

That's unnecessary. A production-grade agent should be able to say:

"I already completed this step. I can reuse the result."

That's where agentic caching becomes important.

1. Tool result cache

1.1 What is it?

Tool calls are one of the biggest caching opportunities in agentic systems.

Suppose an agent has get_customer(), get_orders(), get_inventory(), get_weather(), search_web() and get_exchange_rate(). The agent might call get_customer(customer_id=123) several times during the same task. If the customer information hasn't changed, making the same API call again provides little value.

Instead:

  1. Agent
  2. Tool call
  3. Tool cache
  4. Hitreturn existing resultMissexecute tool
  5. Store result

1.2 Basic implementation

Using Redis:

import json
 
import redis.asyncio as redis
 
redis_client = redis.Redis(host="localhost", port=6379, decode_responses=True)
 
 
async def cached_tool_call(tool_name: str, arguments: dict, tool_fn, ttl: int = 300):
    cache_key = f"tool:{tool_name}:{json.dumps(arguments, sort_keys=True)}"
 
    cached = await redis_client.get(cache_key)
    if cached:
        return json.loads(cached)
 
    result = await tool_fn(**arguments)
 
    await redis_client.setex(cache_key, ttl, json.dumps(result))
    return result

Usage:

customer = await cached_tool_call(
    tool_name="get_customer",
    arguments={"customer_id": "123"},
    tool_fn=get_customer,
    ttl=600,
)

1.3 Production cache key

Don't key on the tool name alone. A tool can be called with different arguments.

Bad:

key = "get_customer"

Better: combine the tool name + arguments + tool version + tenant ID. For example:

import hashlib
import json
 
 
def tool_cache_key(tenant_id: str, tool_name: str, arguments: dict, tool_version: str):
    raw = json.dumps(
        {
            "tenant_id": tenant_id,
            "tool": tool_name,
            "arguments": arguments,
            "version": tool_version,
        },
        sort_keys=True,
    )
    return hashlib.sha256(raw.encode()).hexdigest()

2. Function call cache

2.1 What is it?

Tool caching and function-call caching are closely related, but function-call caching can be applied at a lower level of abstraction. Consider:

  1. LLM
  2. Function call
  3. get_exchange_rate()
  4. API

If the same function arguments produce the same result for a reasonable period, cache the function execution.

2.2 Example

import time
from functools import wraps
 
 
def cached(ttl=300):
    cache = {}
 
    def decorator(fn):
        @wraps(fn)
        def wrapper(*args, **kwargs):
            key = (fn.__name__, args, tuple(sorted(kwargs.items())))
            now = time.time()
 
            if key in cache:
                value, timestamp = cache[key]
                if now - timestamp < ttl:
                    return value
 
            result = fn(*args, **kwargs)
            cache[key] = (result, now)
            return result
 
        return wrapper
 
    return decorator

Usage:

@cached(ttl=60)
def get_exchange_rate(from_currency, to_currency):
    return external_exchange_api(from_currency, to_currency)

Now get_exchange_rate("USD", "INR") doesn't necessarily hit the external API every time.

2.3 Don't cache everything

This is critical. Consider create_payment(), delete_user(), place_order() and transfer_money().

These are side-effecting operations, and blindly caching them can be dangerous. place_order() should not be treated like a read such as get_customer().

For mutations, you usually need idempotency, which we'll cover later.

3. Agent state cache

3.1 What is it?

An agent isn't just producing an answer. It has state. For example:

{
  "task_id": "task_123",
  "goal": "Research company X",
  "current_step": "financial_analysis",
  "completed_steps": ["company_search", "financial_report"],
  "results": {
    "company": "...",
    "revenue": "..."
  }
}

This state allows the agent to resume execution.

3.2 With and without state

  1. Step 1
  2. Step 2
  3. Step 3
  4. Crash
  5. Start again from step 1
Without state: a crash means starting again
  1. Step 1 ✓
  2. Step 2 ✓
  3. Step 3 ✓
  4. Crash
  5. Load state
  6. Step 4
With state: a crash means resuming

3.3 Redis implementation

import json
 
 
async def save_agent_state(task_id: str, state: dict, ttl: int = 3600):
    await redis_client.setex(f"agent:state:{task_id}", ttl, json.dumps(state))
 
 
async def load_agent_state(task_id: str):
    data = await redis_client.get(f"agent:state:{task_id}")
 
    if not data:
        return None
 
    return json.loads(data)

3.4 Agent loop

state = await load_agent_state(task_id)
 
if not state:
    state = {"completed_steps": [], "results": {}}
 
if "search" not in state["completed_steps"]:
    result = await search_company()
 
    state["results"]["company"] = result
    state["completed_steps"].append("search")
 
    await save_agent_state(task_id, state)

The important part is:

Persist state after meaningful checkpoints, not only when the entire task finishes.

4. Workflow and step cache

4.1 What is it?

Agent workflows frequently contain deterministic steps. Consider a research agent:

  1. Search company
  2. Extract revenue
  3. Calculate growth
  4. Find competitors
  5. Generate report

If step 3 has already completed successfully, don't execute it again.

4.2 Step cache architecture

  1. Workflow
  2. Step
  3. Step cache
  4. Hitreturn resultMissexecute step
  5. Store in cache

4.3 Implementation

async def execute_step(workflow_id: str, step_name: str, input_data: dict, step_fn):
    key = make_step_key(workflow_id, step_name, input_data)
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    result = await step_fn(input_data)
 
    await redis_client.setex(key, 3600, json.dumps(result))
    return result

4.4 Why this is powerful

Imagine the workflow fails at step 5 and has to be retried:

Step First run Retry without step cache Retry with step cache
1 5 s 5 s cached
2 8 s 8 s cached
3 12 s 12 s cached
4 20 s 20 s cached
5 10 s 10 s 10 s
Total 55 s 55 s ≈ 10 s

This becomes extremely important for long-running agents.

5. MCP tool cache

5.1 What is it?

Model Context Protocol (MCP) systems introduce another caching opportunity. An agent might reach search, a database, files, a CRM and internal APIs through MCP servers. Repeated MCP calls can add unnecessary latency.

Conceptually:

  1. Agent
  2. MCP client
  3. MCP tool cache
  4. MCP server
  5. External system

5.2 Cache by tool and arguments

def mcp_cache_key(server, tool, arguments, version):
    payload = {
        "server": server,
        "tool": tool,
        "arguments": arguments,
        "version": version,
    }
    raw = json.dumps(payload, sort_keys=True)
    return hashlib.sha256(raw.encode()).hexdigest()

Then:

async def call_mcp_tool(server, tool, arguments):
    key = mcp_cache_key(server, tool, arguments, version="v1")
 
    cached = await redis_client.get(f"mcp:{key}")
    if cached:
        return json.loads(cached)
 
    result = await actual_mcp_call(server, tool, arguments)
 
    await redis_client.setex(f"mcp:{key}", 300, json.dumps(result))
    return result

5.3 Be careful with MCP caching

Not every MCP resource should be cached equally:

Data Cache policy
Static documentation Long TTL
Product metadata Medium TTL
Live inventory Short TTL
Financial transactions Don't blindly cache
User permissions Very careful

The cache policy belongs to the data semantics, not just the tool name.

6. MCP resource cache

6.1 What is it?

MCP isn't only about tools. Resources can also represent data that agents access repeatedly, for example file://company/policy.pdf, db://customers/123 or docs://product/api.

A resource cache avoids fetching the same resource again and again.

6.2 Example

async def get_mcp_resource(resource_uri: str, fetch_fn, ttl: int = 600):
    key = "mcp:resource:" + hashlib.sha256(resource_uri.encode()).hexdigest()
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    resource = await fetch_fn(resource_uri)
 
    await redis_client.setex(key, ttl, json.dumps(resource))
    return resource

6.3 Version-aware resources

If resources have versions, keying on the version too is safer:

  • Any resource: resource URI + resource version
  • Documents: document ID + document version
  • APIs: resource ID + updated_at

7. API response cache

7.1 What is it?

Agents frequently call external services: weather, stock prices, CRM, search, maps. External calls introduce:

  • latency
  • rate limits
  • cost
  • network failures

Caching can reduce all four.

7.2 Production implementation

import json
 
 
async def cached_api_response(key: str, fetch_fn, ttl: int = 300):
    cache_key = f"api:{key}"
 
    cached = await redis_client.get(cache_key)
    if cached:
        return json.loads(cached)
 
    result = await fetch_fn()
 
    await redis_client.setex(cache_key, ttl, json.dumps(result))
    return result

Usage:

weather = await cached_api_response(
    key="weather:bengaluru",
    fetch_fn=lambda: fetch_weather("bengaluru"),
    ttl=300,
)

7.3 Stale-while-revalidate

For some applications, you don't need perfectly fresh data. You can serve a slightly stale result while refreshing the cache in the background:

  1. Request
  2. Cache
  3. FreshreturnStalereturn now, refresh in backgroundMissingfetch

This can dramatically reduce latency.

8. Agent memory cache

8.1 What is it?

Agent memory can exist at multiple levels:

Short-term memory Long-term memory
Current conversation User preferences
Current task Historical facts
Recent tool results Learned context

Not all memory needs to be fetched from a database every time.

8.2 Example

async def get_user_memory(user_id: str):
    key = f"memory:user:{user_id}"
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    memory = await database.get_user_memory(user_id)
 
    await redis_client.setex(key, 1800, json.dumps(memory))
    return memory

Now the agent can access frequently used memory quickly.

8.3 Memory invalidation

Suppose the user changes their preferred language from English to Tamil. Your memory cache must be invalidated:

await redis_client.delete(f"memory:user:{user_id}")

Otherwise the agent may keep using stale information.

9. Planning and reasoning cache

9.1 What is it?

This is one of the more advanced forms of agent caching. Suppose an agent repeatedly receives:

"Analyze this invoice and extract: vendor, total, tax and due date."

The high-level plan may be reusable:

  1. Read invoice
  2. Extract fields
  3. Validate fields
  4. Return structured result

Instead of asking the LLM to generate the same plan every time:

  1. User request
  2. Planning LLM
  3. Plan

you can potentially cache stable plans.

9.2 Architecture

  1. Task
  2. Normalize intent
  3. Planning cache
  4. Hitreturn existing planMisscall planning LLM
  5. Store in cache

9.3 Example

async def get_plan(task):
    normalized = normalize_task(task)
    key = "plan:" + hashlib.sha256(normalized.encode()).hexdigest()
 
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    plan = await planning_model(task)
 
    await redis_client.setex(key, 3600, json.dumps(plan))
    return plan

9.4 Important warning

Reasoning plans are more sensitive than deterministic data. A cached plan may become invalid when:

  • tools change
  • permissions change
  • available data changes
  • system instructions change
  • agent capabilities change

So include these versions in the cache key: task + agent version + toolset version + prompt version.

10. Task and job cache

10.1 What is it?

Long-running agents frequently run asynchronous jobs. For example:

  1. User
  2. "Analyze this 500-page document"
  3. Task ID
  4. Background agent

The user might refresh the page. You don't want to start another job. Instead:

  1. POST /analyze
  2. Task cache
  3. Existing taskreturn task_idNew taskstart agent

10.2 Idempotent task creation

async def create_task(user_id, request):
    request_hash = hashlib.sha256(
        json.dumps(request, sort_keys=True).encode()
    ).hexdigest()
    key = f"task:{user_id}:{request_hash}"
 
    existing = await redis_client.get(key)
    if existing:
        return json.loads(existing)
 
    task = await start_agent_task(user_id, request)
 
    await redis_client.setex(key, 3600, json.dumps(task))
    return task

Now when a user clicks twice, the same request maps to the same task, instead of creating task 1 and task 2.

11. Idempotency cache

11.1 Caching is not idempotency

This deserves special attention. Consider POST /payments.

If the network times out, the client may retry, and the server receives payment request #1 and payment request #2. If both execute, the customer is charged ₹10,000 twice.

11.2 Idempotency key

The client generates a key, for example idempotency_key = abc123, and sends it with the request. The server:

async def process_payment(idempotency_key, payment_data):
    cache_key = f"payment:{idempotency_key}"
 
    existing = await redis_client.get(cache_key)
    if existing:
        return json.loads(existing)
 
    result = await charge_payment(payment_data)
 
    await redis_client.set(cache_key, json.dumps(result), ex=86400)
    return result

Now retries become safe:

  1. Process payment
  2. Store result
Request 1
  1. Look up key
  2. Return previous result
Request 2, same idempotency key

This is essential for agents that can trigger mutations.

12. Agent checkpoint cache

12.1 What is it?

Agent workflows can be extremely long. Consider:

  1. Search
  2. Download
  3. Parse
  4. Analyze
  5. Compare
  6. Validate
  7. Generate report
  8. Send email

If step 7 fails, restarting everything is wasteful. Checkpoint after important steps:

  1. Step 1 ✓checkpoint
  2. Step 2 ✓checkpoint
  3. Step 3 ✓checkpoint
  4. Step 4 ✓checkpoint
  5. Step 5 ✓checkpoint
  6. Step 6 ✓checkpoint
  7. Step 7 ✗
Checkpoint after every completed step

Then resume from the last checkpoint, straight into step 7.

12.2 Implementation

async def checkpoint(task_id, step, state):
    key = f"checkpoint:{task_id}:{step}"
    await redis_client.set(key, json.dumps(state))
 
 
async def load_checkpoint(task_id, step):
    key = f"checkpoint:{task_id}:{step}"
 
    data = await redis_client.get(key)
    if not data:
        return None
 
    return json.loads(data)

13. Sub-agent result cache

13.1 What is it?

Modern agents can delegate tasks:

  1. Main agent
  2. Research agentCoding agentData agentValidation agent

Suppose the research agent has already answered:

"Find the company's revenue for 2025."

The main agent shouldn't necessarily trigger another research agent for the exact same task.

13.2 Architecture

  1. Main agent
  2. Sub-agent request
  3. Sub-agent cache
  4. Hitreturn resultMissrun sub-agent
  5. Store in cache

13.3 Key design

Key on tenant + sub-agent type + task + input + agent version. For example:

def subagent_key(tenant_id, agent_name, task, input_data, version):
    raw = json.dumps(
        {
            "tenant": tenant_id,
            "agent": agent_name,
            "task": task,
            "input": input_data,
            "version": version,
        },
        sort_keys=True,
    )
    return hashlib.sha256(raw.encode()).hexdigest()

14. Multi-agent shared cache

14.1 What is it?

Now consider several agents working together:

  1. Main agent
  2. Research agentData agent
  3. Shared cache

Agents may need to share results. For example:

  1. Research agent
  2. Company revenue = $4.2B
  3. Shared cache
  4. Financial agent

The financial agent doesn't need to fetch the same information again.

14.2 But shared state introduces problems

You need to define:

  • Who can write?
  • Who can read?
  • Who owns the data?
  • How long is it valid?
  • What version produced it?
  • Can another tenant access it?

For enterprise systems, this becomes an authorization problem. A safe key might include tenant ID + namespace + resource ID + version.

15. Long-running agent cache

15.1 What is it?

Some agents run for seconds, some for minutes and some for hours: deep research, data analysis, document processing, code generation, report generation, monitoring.

You don't want the agent's entire state to live only in memory. Instead, the worker loops:

  1. Execute
  2. Checkpoint
  3. Save state
  4. Continue

If the worker dies, another one picks up where it stopped:

  1. Worker 1
  2. Crash
  3. Worker 2
  4. Load checkpoint
  5. Resume

16. Designing a production agent cache

16.1 Putting the pieces together

At this point, we can combine the pieces:

  1. User
  2. API gateway
  3. Task / idempotency cache
  4. Agent state
  5. Agent
  6. PlanningToolsSub-agent
  7. Plan cacheTool cacheSub-agent cache
  8. MCPvia toolsAPIvia tools
  9. MCP cacheAPI cache
  10. Resources
  11. Resource cache

16.2 Cache key design for agents

Cache key design becomes extremely important once multiple agents and tools are involved.

  • A weak key: "get_customer"
  • A better key: tenant ID + tool name + arguments
  • A production-oriented key: tenant ID + user scope + tool name + arguments + tool version + agent version + data version

For example:

def make_agent_cache_key(
    tenant_id,
    user_scope,
    component,
    input_data,
    component_version,
    data_version=None,
):
    payload = {
        "tenant_id": tenant_id,
        "user_scope": user_scope,
        "component": component,
        "input": input_data,
        "component_version": component_version,
        "data_version": data_version,
    }
    raw = json.dumps(payload, sort_keys=True)
    return hashlib.sha256(raw.encode()).hexdigest()

16.3 TTL strategy for agentic systems

Different agent components need different TTLs:

Component Example TTL
Weather tool 5 min
Customer profile 10 min
Search results 5–30 min
Product metadata 1 hour
Agent plan 1 hour
Static documentation 24 hours
Tool result Tool-dependent
MCP resource Resource-dependent
Sub-agent result Task-dependent
Agent state Until task completion
Checkpoint Until task completion
Payment result Idempotency window

These are examples, not universal values. The correct TTL depends on freshness requirements, data volatility, business impact, cost and failure-recovery requirements.

16.4 Cache invalidation in agentic systems

This is where agent caching gets complicated. Suppose the agent reads a customer through a tool backed by a database:

  1. Agent
  2. Tool
  3. Database

In the database, the customer's status changes from ACTIVE to SUSPENDED. But the agent cache still says ACTIVE. Now the agent acts on stale information.

TTL-based

  1. Cache
  2. 10 minutes
  3. Expire

Simple, but not always precise.

Event-based

  1. Database update
  2. Event
  3. Invalidate cache

More accurate.

Version-based

Keep a version on the record, for example customer_version = 42, and put it in the cache key: customer:123:v42. When the customer changes, customer_version becomes 43, and the old cache entry naturally becomes irrelevant.

Explicit invalidation

await redis_client.delete(f"customer:{customer_id}")

Useful for important state changes.

16.5 Caching and agent safety

Caching introduces a new attack surface. Consider:

  1. User A"What is customer 123's balance?"
  2. Response cached
  3. User B"Show me customer 123's balance"
  4. Cache hit
  5. Private data leak

If the cache isn't authorization-aware, User B sees data they may not be allowed to see. So the cache key should combine tenant + user permissions + resource + query + version.

16.6 Never cache sensitive results blindly

Be particularly careful with:

  • passwords
  • authentication and access tokens
  • payment information
  • private customer information
  • personal data
  • healthcare information
  • financial information
  • secrets

Caching can turn a normal application bug into a cross-user data leak.

16.7 Cache stampede

Another important production problem is the cache stampede. Imagine a popular cache entry expires:

  1. Cache expires
  2. 10,000 requests
  3. 10,000 API calls

You just destroyed the purpose of caching.

The fix is a lock, so only one worker refreshes the entry:

import asyncio
import json
 
 
async def get_with_lock(key, fetch_fn, ttl=300):
    cached = await redis_client.get(key)
    if cached:
        return json.loads(cached)
 
    lock_key = f"{key}:lock"
    acquired = await redis_client.set(lock_key, "1", nx=True, ex=30)
 
    if acquired:
        try:
            result = await fetch_fn()
            await redis_client.setex(key, ttl, json.dumps(result))
            return result
        finally:
            await redis_client.delete(lock_key)
 
    # Another worker is refreshing. Wait briefly and retry.
    await asyncio.sleep(0.1)
    return await get_with_lock(key, fetch_fn, ttl)

This prevents thousands of workers from performing the same expensive operation.

16.8 Cache observability

Production agents need cache telemetry. At minimum, track cache hits, misses, errors, evictions, latency, size, stale reads and refresh count.

For agents, also track tool calls avoided, agent steps avoided, LLM calls avoided, API calls avoided, tokens saved, agent runtime saved and cost saved.

An agent cache dashboard might show:

Metric Value
Tool cache hit rate 72.4%
MCP cache hit rate 61.8%
Agent state hit rate 98.1%
Plan cache hit rate 84.3%
API calls avoided 42K
Agent steps avoided 18K
LLM calls avoided 9K
Tokens saved 14.2M
Estimated cost saved $980

17. Cache, state and idempotency

17.1 Caching vs idempotency

These concepts are related but different.

  • Caching: "I already have the answer."
  • Idempotency: "I already processed this operation."

GET customer is a good fit for caching. POST payment requires idempotency.

This distinction is extremely important when building agents that can take actions.

17.2 Caching vs state

Another important distinction:

  • Cache: temporary, reusable data, such as tool results, search results and API responses.
  • State: information required to continue execution, such as the current task, completed steps, the agent plan, checkpoints and workflow status.

State may use the same infrastructure as a cache, such as Redis, but conceptually they are different. For critical state, you may want durable storage in addition to Redis:

  1. Agent
  2. Redisfast stateDatabasedurable state

17.3 A shared cache layer

Instead of letting every component (tools, MCP, the agent, API clients) talk to Redis directly and implement caching differently, create a common abstraction:

import json
 
 
class Cache:
    def __init__(self, redis_client):
        self.redis = redis_client
 
    async def get(self, key):
        value = await self.redis.get(key)
        if value is None:
            return None
        return json.loads(value)
 
    async def set(self, key, value, ttl):
        await self.redis.setex(key, ttl, json.dumps(value))
 
    async def delete(self, key):
        await self.redis.delete(key)

Now cache = Cache(redis_client) can be shared across the agent, tool layer, MCP layer, workflow layer, API layer and memory layer.

For a serious production system:

  1. Agent
  2. PlanningToolsWorkflow
  3. Plan cacheTool cacheStep cache
  4. MCP cachevia toolsAPI cachevia tools
  5. Agent state
  6. Redisfast storePostgreSQLdurable store

This gives you fast access, recovery, deduplication, lower cost, lower latency and better scalability.

18. Before you ship

18.1 Production checklist

Before deploying agent caching, ask:

Tool caching

  • Is this tool safe to cache?
  • What is the TTL?
  • Are arguments included in the key?
  • Is the tool version included?
  • Is tenant isolation enforced?

Agent state

  • Can the agent resume after failure?
  • Are checkpoints persisted?
  • Is state durable?
  • Can two workers modify the same state?

MCP

  • Which tools are cacheable?
  • Which resources are cacheable?
  • Are resource versions tracked?
  • Are permissions included?

APIs

  • Is the API response deterministic?
  • Does the provider allow caching?
  • What is the freshness requirement?
  • What happens when the cache is stale?

Idempotency

  • Are mutations protected?
  • Are retries safe?
  • Is there an idempotency key?
  • Can duplicate agent actions happen?

Operations

  • Cache hit and miss metrics?
  • Cache latency?
  • Cache size?
  • Eviction policy?
  • Stampede protection?
  • Invalidation strategy?
  • Security audit?

19. The full picture

19.1 The core principle

Agentic caching isn't about putting Redis between your agent and everything else. It's about identifying repeated work. An agent may repeatedly perform the same:

  1. Plan
  2. Tool call
  3. API request
  4. MCP resource fetch
  5. Workflow step
  6. Sub-agent task

A production system should recognize "I have already done this" and safely reuse the result.

19.2 Part 2 takeaway

The major caching layers for agentic AI are:

  1. Tool result cache
  2. Function call cache
  3. Agent state cache
  4. Workflow / step cache
  5. MCP tool cache
  6. MCP resource cache
  7. API response cache
  8. Agent memory cache
  9. Planning / reasoning cache
  10. Task / job cache
  11. Idempotency cache
  12. Agent checkpoint cache
  13. Sub-agent result cache
  14. Multi-agent shared cache
  15. Long-running agent cache

The most important production distinction is:

Concept What it does
Cache Reuse previous work
State Resume previous work
Idempotency Prevent duplicate work

And the overall architecture becomes:

  1. User
  2. Agent
  3. PlanningToolsWorkflow
  4. Plan cacheTool cacheStep cache
  5. MCPvia toolsAPIsvia tools
  6. MCP cacheAPI cache
  7. Resources
  8. Agent state
  9. RedisDurable DB

The best agent isn't the one that calls the most tools. It's the one that knows when it doesn't need to call them again.

Next: Part 3, Conversational AI caching covers context, sessions, memory and responses.