HaribaskarAI Engineer
← All posts

The AI Caching Playbook, Part 5: Cache the Agent's Work, Not Just the Answer

Agent state, plans, tool and API results, workflow steps, checkpoints, idempotent writes and scoped sharing: caching the work an agent does.

Haribaskar Dhanabalan19 min read

A glowing cache server in a dark data centre, with circuit lines connecting an AI node and other services to it
On this page

This is Part 5 of the AI Caching Playbook. The earlier parts looked at caching in LLM applications, agents, conversational systems and RAG pipelines.

But agentic AI introduces a much harder problem. A traditional RAG system might do:

  1. User
  2. Retrieve
  3. Rerank
  4. LLM
  5. Answer

An agent can do this:

  1. User
  2. Agent
  3. Plan
  4. Search
  5. Tool
  6. API
  7. Database
  8. Another tool
  9. Reason
  10. Search again
  11. Call another agent
  12. Reason
  13. Final answer

The agent isn't performing one expensive operation. It is performing a sequence of operations, and every one of them can potentially be repeated.

That changes the caching problem completely.

1. The core problem with agentic AI

1.1 One request, many steps

Consider an agent that receives:

"Find the best laptop under ₹80,000, compare three options, check their current prices, and recommend one."

The agent might execute:

  1. Understand the request
  2. Create a plan
  3. Search products
  4. Search reviews
  5. Fetch product details
  6. Check prices
  7. Compare specifications
  8. Rank products
  9. Generate a recommendation

Without caching, every request repeats everything. With caching, every request reuses previous work and executes only what changed:

  1. Every request
  2. Reuse previous work
  3. Execute only what changed
With caching

The goal becomes:

Cache the work performed by the agent, not just the final response.

1.2 Agentic caching architecture

A production agent can have caching at many levels:

  1. User
  2. Response cacheon a miss, continue
  3. Agent state cache
  4. Planning cache
  5. Agent execution
  6. Tool cacheAPI cacheDB cache
  7. Intermediate results
  8. Agent context cache
  9. Checkpoint / recovery
  10. Answer

Each layer has a different purpose.

2. Agent state cache

2.1 What is it?

The first thing to cache is the agent's current state. For example:

{
  "goal": "Find the best laptop",
  "status": "researching",
  "current_step": 3,
  "products_found": 12,
  "selected_products": 5
}

Instead of reconstructing this state from the entire conversation every time (conversation → LLM → infer the current state), store it explicitly and give the agent both:

  1. Conversation + explicit state
  2. Agent

2.2 Redis implementation

import json
 
import redis
 
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
 
 
def state_key(tenant_id, user_id, session_id):
    return f"agent:state:{tenant_id}:{user_id}:{session_id}"
 
 
def save_agent_state(tenant_id, user_id, session_id, state):
    key = state_key(tenant_id, user_id, session_id)
    r.setex(key, 3600, json.dumps(state))
 
 
def get_agent_state(tenant_id, user_id, session_id):
    key = state_key(tenant_id, user_id, session_id)
 
    value = r.get(key)
    if not value:
        return None
 
    return json.loads(value)

This gives the agent a durable working state.

2.3 Why explicit state matters

Without explicit state, when the agent asks itself "What have I already checked?", the LLM has to reconstruct the answer from context.

With explicit state, the agent already knows:

state.current_step = 4
state.completed_tools = [...]
state.results = [...]

This reduces tokens, latency, repeated reasoning, context size and cost. A useful principle:

Don't make the LLM remember state that your application can represent deterministically.

3. Planning cache

3.1 What is it?

Agents often generate plans. For the goal "Research three laptops", the plan might be:

  1. Search products
  2. Get specifications
  3. Get prices
  4. Compare
  5. Recommend

Generating the plan may itself require an LLM call. If the same type of task appears repeatedly, the plan can sometimes be cached, keyed on goal + agent type + planning prompt version + model version.

3.2 Cache key

import hashlib
import json
 
 
def planning_cache_key(goal, agent_type, model_version, prompt_version):
    payload = {
        "goal": " ".join(goal.lower().split()),
        "agent_type": agent_type,
        "model": model_version,
        "prompt": prompt_version,
    }
    digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    return f"agent:plan:{digest}"

Then:

def get_or_create_plan(goal, agent_type, planner):
    key = planning_cache_key(goal, agent_type, "model-v3", "prompt-v4")
 
    cached = r.get(key)
    if cached:
        return json.loads(cached)
 
    plan = planner.create_plan(goal)
 
    r.setex(key, 1800, json.dumps(plan))
    return plan

3.3 But don't cache every plan

Planning is context-sensitive. "Book a flight to Chennai" can depend on the current date, budget, user preferences, passport requirements and available flights.

So plan caching is safest when:

  • the task is deterministic
  • the plan doesn't depend on volatile state
  • the environment hasn't changed
  • the same tool set is available

For dynamic agents, cache reusable plan templates rather than blindly reusing complete plans. For example, a research task always follows the same shape, while the individual values stay dynamic:

  1. Search
  2. Compare
  3. Validate
  4. Summarize
A reusable plan template

4. Tool result cache

4.1 What is it?

This is one of the most valuable caches in agentic systems. Suppose an agent calls:

weather_tool(city="Bangalore")

and receives:

{
  "temperature": 27,
  "condition": "Cloudy"
}

If another request asks for Bangalore's weather 30 seconds later, you may not need another external API call:

  1. Agent
  2. Tool
  3. Cache
  4. Hitreturn resultMisscall the external API
  5. Store result in cache

4.2 A generic tool cache

Instead of implementing caching separately for every tool, create a wrapper:

import hashlib
import json
 
 
def tool_cache_key(tool_name, arguments, tool_version):
    payload = {
        "tool": tool_name,
        "version": tool_version,
        "arguments": arguments,
    }
    digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    return f"agent:tool:{digest}"

Now:

def cached_tool_call(tool_name, arguments, tool_version, tool_function, ttl=300):
    key = tool_cache_key(tool_name, arguments, tool_version)
 
    cached = r.get(key)
    if cached:
        return json.loads(cached)
 
    result = tool_function(**arguments)
 
    r.setex(key, ttl, json.dumps(result))
    return result

This creates a reusable tool caching layer.

4.3 Deterministic vs non-deterministic tools

Not every tool should be cached the same way:

Category Examples
Highly cacheable Currency conversion, historical weather, product specifications, company documentation, static metadata, public FAQ, geocoding
Short TTL Current weather, stock prices, flight availability, hotel availability, inventory, exchange rates
Usually don't cache blindly Create payment, place order, send email, delete account, transfer money, submit form

The last group are side-effecting operations. Caching them can be dangerous.

4.4 Read tools vs write tools

A useful production classification:

Tool type Example Treatment
Read get_customer() Can potentially be cached
Write create_customer() Needs idempotency and explicit execution semantics

You should not run cache.get(key) before a write operation and assume that prevents duplicate execution. Instead, use an idempotency key + an execution record + transaction semantics.

5. API response cache

5.1 What is it?

Agents frequently call APIs: weather, search, maps, CRM. Every external API call adds latency, cost, rate-limit pressure and failure risk. A generic API cache can help:

def api_cache_key(service, endpoint, params, version="v1"):
    payload = {
        "service": service,
        "endpoint": endpoint,
        "params": params,
        "version": version,
    }
    digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    return f"api:{digest}"

Then:

def cached_api_request(service, endpoint, params, request_function, ttl=60):
    key = api_cache_key(service, endpoint, params)
 
    cached = r.get(key)
    if cached:
        return json.loads(cached)
 
    response = request_function(endpoint, params)
 
    r.setex(key, ttl, json.dumps(response))
    return response

6. Database query cache

6.1 What is it?

Agents may also query databases repeatedly (agent → SQL → database → result). A safe, read-only query can sometimes be cached:

def db_query_cache_key(database, query, parameters, schema_version):
    payload = {
        "database": database,
        "query": query,
        "parameters": parameters,
        "schema": schema_version,
    }
    digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    return f"db:query:{digest}"

Then:

def cached_query(query, parameters, execute, ttl=30):
    key = db_query_cache_key("production", query, parameters, "schema-v8")
 
    cached = r.get(key)
    if cached:
        return json.loads(cached)
 
    result = execute(query, parameters)
 
    r.setex(key, ttl, json.dumps(result))
    return result

But database caching needs a careful invalidation strategy. After an UPDATE customer, the related query cache may be stale.

6.2 Cache database results by data version

A stronger approach is to keep a data version, for example customer_data_version = 481. When the relevant data changes, the version moves from 481 to 482, and because the version is part of the key (query + data version), the cache key changes with it.

This is often easier to reason about than trying to find every query affected by an update.

7. Function result cache

7.1 What is it?

Not every agent tool calls an external API. You may have expensive internal functions such as calculate_risk(), calculate_shipping(), generate_report(), classify_document() and score_customer(). These can be cached too:

def cached_function(function_name, function_version, arguments, function, ttl=600):
    key = tool_cache_key(function_name, arguments, function_version)
 
    cached = r.get(key)
    if cached:
        return json.loads(cached)
 
    result = function(**arguments)
 
    r.setex(key, ttl, json.dumps(result))
    return result

This turns caching into an infrastructure capability rather than an application-specific feature.

8. Agent trajectory cache

8.1 What is it?

An agent's execution can be represented as a trajectory:

  1. User
  2. Plan
  3. Tool A
  4. Tool B
  5. Tool C
  6. Final reasoning

We can store this trajectory. For example:

{
  "goal": "Research product",
  "steps": [
    { "tool": "search", "input": "product A", "output": "..." },
    { "tool": "pricing", "input": "product A", "output": "..." }
  ]
}

This is useful for debugging, replay, evaluation, recovery, auditing and optimization.

But trajectory reuse must be handled carefully. A previous trajectory is not automatically valid in a new environment.

8.2 Reusing agent trajectories

Suppose an agent repeatedly receives:

"Generate our weekly sales report."

The workflow may be:

  1. Query the sales DB
  2. Query the returns DB
  3. Calculate metrics
  4. Generate a chart
  5. Summarize

Instead of letting the LLM rediscover this workflow every time, map the goal to the known workflow and run it on the current data:

  1. Goal
  2. Known workflow
  3. Execute on current data

This is much more reliable. Notice what we're caching: the workflow, not the old data. That's an important distinction.

9. Workflow step and partial execution cache

9.1 Workflow step cache

A better abstraction is to cache individual workflow steps. For the weekly report, if only step 1's input changes:

Step Action
1. Fetch sales Recompute
2. Fetch returns Reuse
3. Calculate metrics Reuse
4. Generate chart Reuse
5. Summarize Reuse

This is much better than rerunning the entire agent.

9.2 Partial execution cache

This is one of the most powerful concepts in agentic caching. Suppose a task runs steps A → B → C → D → E, and execution fails at D.

  1. A
  2. B
  3. C
  4. D
Without checkpoints: restart from the beginning
  1. A ✓
  2. B ✓
  3. C ✓
  4. Resume from D
With checkpoints: resume where it failed

This can save enormous amounts of computation.

10. Agent checkpointing

10.1 Implementation

Store the execution state after important steps:

def checkpoint_key(execution_id, step_id):
    return f"agent:checkpoint:{execution_id}:{step_id}"
 
 
def save_checkpoint(execution_id, step_id, state):
    key = checkpoint_key(execution_id, step_id)
    r.set(key, json.dumps(state))
 
 
def load_checkpoint(execution_id, step_id):
    key = checkpoint_key(execution_id, step_id)
 
    value = r.get(key)
    if not value:
        return None
 
    return json.loads(value)

Now a long-running agent can resume execution.

10.2 Checkpoint vs cache

These are related but not identical:

Goal Example
Cache Avoid recomputation "Search result for query X"
Checkpoint Recover execution state "Agent execution 123 reached step 4 with these intermediate results"

A production agent often needs both.

11. Agent context cache

11.1 What is it?

Agent context can become enormous: the system prompt, the conversation, the plan, tool outputs, RAG results and previous reasoning state. Instead of reconstructing everything, maintain structured context:

agent_context = {
    "goal": "...",
    "plan": [...],
    "state": {...},
    "tool_results": {...},
    "retrieved_docs": [...],
    "recent_events": [...],
}

Then cache the context object:

def agent_context_key(session_id, context_version):
    return f"agent:context:{session_id}:{context_version}"

11.2 Context versioning

Suppose the agent context is at version 12. After a tool result arrives it becomes version 13, and after another tool call, version 14:

  1. Context v12
  2. Tool resultv13
  3. Another tool callv14

The cache key becomes agent:context:session123:v14. This is safer than mutating a single cache entry without tracking versions.

12. Agent memory cache

12.1 Short-term state and long-term memory

Long-running agents may have user preferences, past tasks, important facts, previous outcomes and learned patterns. Store durable memory separately from short-term execution state:

  1. Agent
  2. Short-term stateRedisLong-term memoryDB / vector store

Don't treat all memory as the same thing:

Short-term Long-term
Holds Current task, plan, tool results, conversation and execution state User preferences, historical facts, past interactions, important persistent knowledge
Usually lives in Redis A database, vector store or other persistent storage

Caching can speed up access to long-term memory without making Redis the source of truth.

13. Semantic tool cache

13.1 What is it?

Exact tool caching handles search("latest AI news"). A semantic tool cache might also treat "recent AI news" as the same request.

However, this should only be used when semantic equivalence is safe. For a search tool, "What happened in AI today?" and "What happened in AI yesterday?" are semantically close but need different results.

So semantic tool caching needs semantic similarity + time constraints + tool semantics + freshness.

14. Cache-aware tools

14.1 Cache-aware tool design

A powerful production pattern is to make tools explicitly cache-aware. Instead of exposing only search(query), define search(query, cache_policy), for example:

cache_policy = {
    "enabled": True,
    "ttl": 300,
    "scope": "tenant",
    "freshness": "5m",
}

The agent runtime can then apply consistent caching rules.

14.2 Tool metadata should describe cacheability

A tool registry could contain:

tools = {
    "weather": {
        "cacheable": True,
        "ttl": 300,
        "side_effect": False,
    },
    "get_customer": {
        "cacheable": True,
        "ttl": 30,
        "side_effect": False,
    },
    "send_email": {
        "cacheable": False,
        "side_effect": True,
    },
}

Now the agent runtime knows how to treat each tool. This is much safer than letting every agent developer invent their own rules.

15. Side effects: idempotency and retries

15.1 Idempotency for write tools

Caching isn't enough for side-effecting operations. Suppose the agent calls send_email() and the request times out. The agent doesn't know whether the email was sent, so it retries, and two emails may go out.

The right solution is often idempotency, not caching:

idempotency_key = f"{execution_id}:{tool_name}:{step_id}"

The server stores the idempotency key with the execution result. Retries then return the original result instead of executing the operation again.

15.2 Retry and cache interaction

When a tool call hits an API timeout, before retrying the system should determine whether the operation was:

  • read-only
  • partially completed
  • side-effecting
  • idempotent

For reads, a plain retry is usually straightforward. For writes, retry + idempotency is much safer.

16. Multi-agent cache sharing

16.1 Sharing work between agents

Now imagine a supervisor with three agents that all need the same information:

  1. Supervisor agent
  2. Research agentData agentWriter agent
  3. Shared tool cache

Without shared caching, each agent calls the API itself. With a shared tool cache, the first call is reused by the others. This can significantly reduce duplicated work.

But shared caches need strict scope.

16.2 What can be shared across agents?

Potentially shareable Usually not blindly shareable
Public search results User-specific data
Public documentation Private customer records
Product metadata Authorization-sensitive results
Static configuration Personalized recommendations
Common embeddings Private tool outputs
Common reference data

The cache scope should be explicit: global, tenant, team, user, session or execution.

17. Agent response cache

17.1 What is it?

At the top level, we can still cache final responses. But the response cache key should include the agent's important dependencies:

def agent_response_key(
    tenant_id,
    user_scope,
    task,
    agent_version,
    toolset_version,
    context_version,
):
    payload = {
        "tenant": tenant_id,
        "scope": user_scope,
        "task": task,
        "agent": agent_version,
        "tools": toolset_version,
        "context": context_version,
    }
    digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    return f"agent:response:{digest}"

17.2 Why agent response caching is harder

An agent response can depend on the user, the time, tools, external APIs, the database, memory, RAG, the agent version, tool versions and execution state. So for agents especially:

Same prompt ≠ same answer.

"What's the weather?" should almost never have a long-lived response cache. But "What are the steps in our internal onboarding process?" may be highly cacheable.

18. Dependencies, versions and invalidation

18.1 The agent cache dependency graph

A production agent can have a dependency graph like:

  1. User request
  2. Agent plan
  3. Tool ATool BTool C
  4. Result AResult BResult C
  5. Agent context
  6. Next step
  7. Final answer

If Tool B's result changes, everything downstream of it may need recomputing:

  1. Tool B
  2. Result B
  3. Context
  4. Next step
  5. Final answer

This is why dependency tracking becomes more important as agents become more complex.

18.2 Cache invalidation in agents

There is no universal TTL that solves agent caching. Instead, ask:

What caused this result to become invalid?

Cause Version change
Tool version changed tool-v1 → tool-v2
Agent prompt changed prompt-v4 → prompt-v5
External data changed data-v20 → data-v21
User state changed cart-v8 → cart-v9
Permissions changed role-v2 → role-v3

These versions can be part of cache keys.

18.3 Version everything important

A useful production pattern is to version the agent, prompts, tools, workflows, knowledge, schema and data. Then:

cache key = input + agent version + prompt version + tool version + data version

Changing a version naturally creates a new cache namespace.

19. Long-running agents and recovery

19.1 Long-running agents

Some agents don't finish in seconds. Research agents, data analysis agents, report generators, software engineering agents and workflow automation agents may run for minutes or hours.

For them, caching is less about response reuse and more about execution recovery. This is where checkpoints become critical.

19.2 Production agent execution

Instead of:

result = run_agent(task)

think:

execution = start_execution(task)
 
while not execution.complete:
    step = execution.next_step()
    result = execute_step(step)
 
    save_checkpoint(execution.id, step.id, execution.state)

Now if the process crashes, it restarts, loads the checkpoint and resumes, rather than starting from zero.

19.3 Failure recovery

Imagine steps 1–3 succeeded, step 4 failed and step 5 never started. Recovery should look like:

  1. Load checkpoint
  2. Retry step 4
  3. Step 5
  4. Final response

not a rerun of steps 1 through 5. This is one of the biggest benefits of caching intermediate execution state.

20. Agent caching meets RAG caching

20.1 Agentic RAG

Agentic RAG combines both systems:

  1. Agent
  2. Query transformation cache
  3. Embedding cache
  4. Retrieval cache
  5. Reranking cache
  6. Agent tool cache
  7. Agent state
  8. Agent context
  9. LLM

Now caching happens across several dimensions:

Area Caches
RAG Embedding cache, retrieval cache, reranking cache
Tools API cache, DB cache, function cache
State Session state, execution state
Memory Short-term, long-term

This is where a unified cache architecture becomes valuable.

21. A shared agent cache layer

21.1 A generic agent cache wrapper

We can create one reusable abstraction:

import json
 
 
class AgentCache:
    def __init__(self, redis_client):
        self.redis = redis_client
 
    def get(self, key):
        value = self.redis.get(key)
        if not value:
            return None
        return json.loads(value)
 
    def set(self, key, value, ttl):
        self.redis.setex(key, ttl, json.dumps(value))
 
    def get_or_compute(self, key, compute, ttl=300):
        cached = self.get(key)
        if cached is not None:
            return cached
 
        result = compute()
        self.set(key, result, ttl)
        return result

Usage:

cache = AgentCache(r)
 
result = cache.get_or_compute(
    key="agent:tool:weather:bangalore",
    compute=lambda: weather_tool("Bangalore"),
    ttl=300,
)

This creates a consistent caching interface across the agent runtime.

21.2 Add a cache policy

A more production-ready version accepts a policy:

cache_policy = {
    "enabled": True,
    "ttl": 300,
    "scope": "tenant",
    "stale_while_revalidate": True,
    "semantic": False,
}

The runtime can then decide:

  • Should I cache?
  • For how long?
  • Who can reuse it?
  • Can stale data be returned?
  • Can semantic matching be used?

This moves caching from ad-hoc code into an explicit system design.

22. Observability

22.1 Observability for agent caching

Traditional cache metrics aren't enough. Track agent.cache.hit, agent.cache.miss, agent.cache.stale and agent.cache.lock_wait, but also:

  • agent.tool.calls_avoided
  • agent.llm.calls_avoided
  • agent.steps_reused
  • agent.steps_recomputed
  • agent.tokens_saved
  • agent.execution_time_saved
  • agent.cost_saved

For example:

metrics.increment("agent.tool.cache_hit", tags={"tool": tool_name})
metrics.increment("agent.execution.step_reused")

22.2 Measure agent-level savings

Suppose a typical agent run looks like this before and after caching:

Metric Before caching After caching
Tool calls 12 4
LLM calls 8 3
Latency 25 s 9 s
Cost $0.40 $0.12

The important metric isn't just "cache hit rate = 70%". The real business metrics are cost, from $0.40 to $0.12, and latency, from 25 s to 9 s. That's what matters.

22.3 Don't optimize for cache hit rate alone

A 99% hit rate can still be bad. A 99% hit rate on a cheap operation may be worth far less than a 10% hit rate on an expensive LLM operation.

A better metric:

Cache value = avoided computation cost × reuse frequency × correctness confidence

23. Security and scope

23.1 Security considerations

Agent caching introduces significant security risks. Never accidentally share:

  • user secrets
  • API responses
  • private documents
  • authorization-sensitive results
  • customer records
  • tool credentials
  • private agent memory

And never cache credentials (API keys, passwords, OAuth tokens, session tokens) as if they were normal tool results.

23.2 Cache scope

Every cache should answer:

Who is allowed to reuse this result?

The possible scopes are global, tenant, team, user, session and execution. For example:

Data Scope
Public documentation Global
Company pricing Tenant
User preferences User
Current workflow state Execution

This makes cache boundaries explicit.

24. Before you ship

24.1 Production agent caching checklist

State

  • Agent state is explicit
  • Session state is cached
  • Execution state is versioned
  • Checkpoints exist for long-running tasks

Planning

  • Planning cache evaluated
  • Agent version included
  • Prompt version included
  • Dynamic plans aren't blindly reused

Tools

  • Tools classified as read or write
  • Cacheability defined
  • TTL defined per tool
  • Tool version included

APIs

  • API response caching evaluated
  • Freshness requirements defined
  • Rate limits considered
  • Errors aren't cached blindly

Databases

  • Read queries cached where appropriate
  • Data versioning considered
  • Mutation invalidation defined

Execution

  • Workflow steps are independently cacheable
  • Partial results can be reused
  • Checkpointing implemented
  • Failed executions can resume

Multi-agent

  • Shared cache boundaries defined
  • Tenant isolation enforced
  • User-specific results protected

Security

  • Authorization scope included
  • Sensitive data protected
  • Credentials never cached as ordinary results
  • Cache access monitored

Observability

  • Hit rate
  • Miss rate
  • Steps reused
  • Tool calls avoided
  • LLM calls avoided
  • Tokens saved
  • Cost saved
  • Latency saved

25. The full picture

25.1 The complete agentic cache architecture

Putting everything together:

  1. User
  2. Response cacheon a miss, continue
  3. Agent state cache
  4. Planning cache
  5. Agent loop
  6. Tool cacheRAG cacheAPI cache
  7. Function cacheRetrieval cacheExternal API
  8. Intermediate results
  9. Agent context cache
  10. Checkpoint storecontinue the loop until done
  11. Final answer

25.2 The big shift in thinking

Traditional caching asks:

"Can I cache the response?"

RAG caching asks:

"Can I cache the retrieval?"

Agentic caching asks a much bigger question:

"Which parts of the agent's work can I safely reuse?"

That could be the plan, a tool call, an API result, a database result, a search result, an embedding, a reranking, an intermediate calculation, a workflow step, agent state, agent memory, context, a checkpoint or the final response.

The agent doesn't always need to start from zero.

25.3 The production principle

A naive agent does everything on every request:

  1. Every request
  2. Think
  3. Plan
  4. Call tools
  5. Retrieve
  6. Compute
  7. Think again
  8. Answer

A production agent does this instead:

  1. Every request
  2. What have I already computed?
  3. What is still valid?
  4. What has changed?
  5. Reuse valid work
  6. Compute only the delta
  7. Continue execution

That's the real value of agentic caching.

Don't cache just the answer. Cache the work that produced the answer.

25.4 The final mental model

Think of an agent execution as a computation graph:

  1. User request
  2. Plan
  3. SearchAPIDB
  4. Result AResult BResult C
  5. Analysis
  6. Decision
  7. Tool
  8. Answer

Each node is potentially a cacheable computation:

If… Then…
The search has already run and is still valid Reuse it
The API data has changed Recompute it
The DB result is user-specific Scope it correctly
The tool has side effects Use idempotency
The execution fails Resume from the checkpoint

That is how caching evolves from a simple optimization into an execution architecture.

25.5 Part 5 core takeaways

If you remember only ten things:

  1. Cache agent state.
  2. Cache deterministic tool results.
  3. Separate read tools from write tools.
  4. Use idempotency for side effects.
  5. Cache workflow steps, not only final responses.
  6. Checkpoint long-running executions.
  7. Version agents, prompts, tools and data.
  8. Make cache scope explicit.
  9. Track work avoided, not just cache hits.
  10. Design agents to reuse previous computation.

The ultimate goal is simple:

A production agent should not redo work that is still valid.

And that leads to the next level of the caching architecture.

What's next?

Part 6: LLM inference caching

We've now covered caching around the application: conversation, RAG, agents and tools. But there is another layer where enormous amounts of computation happen, inside the model:

  1. LLM
  2. PrefillDecode
  3. KV cacheTokens

Part 6 will go deeper into:

  • KV cache, prefix cache, prompt caching and prefill caching
  • decode-time caching, attention KV reuse and session-level KV reuse
  • prefix matching, cache-friendly prompt design, and static vs dynamic prompt sections
  • long-context caching, multi-turn KV reuse and batch inference caching
  • speculative decoding and continuous batching
  • GPU memory constraints, KV cache quantization and KV cache eviction
  • paged attention and prefix sharing
  • inference cost optimization, latency vs memory trade-offs, and production LLM serving architecture

The central question becomes:

What if we could avoid making the model recompute the same attention work every time?

That takes us from application-level caching into the internals of LLM inference itself.

Next: Part 6, LLM inference caching covers the KV cache, prefix caching and prompt caching inside the model.