The AI Caching Playbook, Part 5: Cache the Agent's Work, Not Just the Answer
Agent state, plans, tool and API results, workflow steps, checkpoints, idempotent writes and scoped sharing: caching the work an agent does.

On this page
This is Part 5 of the AI Caching Playbook. The earlier parts looked at caching in LLM applications, agents, conversational systems and RAG pipelines.
But agentic AI introduces a much harder problem. A traditional RAG system might do:
- User
- Retrieve
- Rerank
- LLM
- Answer
An agent can do this:
- User
- Agent
- Plan
- Search
- Tool
- API
- Database
- Another tool
- Reason
- Search again
- Call another agent
- Reason
- Final answer
The agent isn't performing one expensive operation. It is performing a sequence of operations, and every one of them can potentially be repeated.
That changes the caching problem completely.
1. The core problem with agentic AI
1.1 One request, many steps
Consider an agent that receives:
"Find the best laptop under ₹80,000, compare three options, check their current prices, and recommend one."
The agent might execute:
- Understand the request
- Create a plan
- Search products
- Search reviews
- Fetch product details
- Check prices
- Compare specifications
- Rank products
- Generate a recommendation
Without caching, every request repeats everything. With caching, every request reuses previous work and executes only what changed:
- Every request
- Reuse previous work
- Execute only what changed
The goal becomes:
Cache the work performed by the agent, not just the final response.
1.2 Agentic caching architecture
A production agent can have caching at many levels:
- User
- Response cacheon a miss, continue
- Agent state cache
- Planning cache
- Agent execution
- Tool cacheAPI cacheDB cache
- Intermediate results
- Agent context cache
- Checkpoint / recovery
- Answer
Each layer has a different purpose.
2. Agent state cache
2.1 What is it?
The first thing to cache is the agent's current state. For example:
{
"goal": "Find the best laptop",
"status": "researching",
"current_step": 3,
"products_found": 12,
"selected_products": 5
}Instead of reconstructing this state from the entire conversation every time (conversation → LLM → infer the current state), store it explicitly and give the agent both:
- Conversation + explicit state
- Agent
2.2 Redis implementation
import json
import redis
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
def state_key(tenant_id, user_id, session_id):
return f"agent:state:{tenant_id}:{user_id}:{session_id}"
def save_agent_state(tenant_id, user_id, session_id, state):
key = state_key(tenant_id, user_id, session_id)
r.setex(key, 3600, json.dumps(state))
def get_agent_state(tenant_id, user_id, session_id):
key = state_key(tenant_id, user_id, session_id)
value = r.get(key)
if not value:
return None
return json.loads(value)This gives the agent a durable working state.
2.3 Why explicit state matters
Without explicit state, when the agent asks itself "What have I already checked?", the LLM has to reconstruct the answer from context.
With explicit state, the agent already knows:
state.current_step = 4
state.completed_tools = [...]
state.results = [...]This reduces tokens, latency, repeated reasoning, context size and cost. A useful principle:
Don't make the LLM remember state that your application can represent deterministically.
3. Planning cache
3.1 What is it?
Agents often generate plans. For the goal "Research three laptops", the plan might be:
- Search products
- Get specifications
- Get prices
- Compare
- Recommend
Generating the plan may itself require an LLM call. If the same type of task appears repeatedly, the plan can sometimes be cached, keyed on goal + agent type + planning prompt version + model version.
3.2 Cache key
import hashlib
import json
def planning_cache_key(goal, agent_type, model_version, prompt_version):
payload = {
"goal": " ".join(goal.lower().split()),
"agent_type": agent_type,
"model": model_version,
"prompt": prompt_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"agent:plan:{digest}"Then:
def get_or_create_plan(goal, agent_type, planner):
key = planning_cache_key(goal, agent_type, "model-v3", "prompt-v4")
cached = r.get(key)
if cached:
return json.loads(cached)
plan = planner.create_plan(goal)
r.setex(key, 1800, json.dumps(plan))
return plan3.3 But don't cache every plan
Planning is context-sensitive. "Book a flight to Chennai" can depend on the current date, budget, user preferences, passport requirements and available flights.
So plan caching is safest when:
- the task is deterministic
- the plan doesn't depend on volatile state
- the environment hasn't changed
- the same tool set is available
For dynamic agents, cache reusable plan templates rather than blindly reusing complete plans. For example, a research task always follows the same shape, while the individual values stay dynamic:
- Search
- Compare
- Validate
- Summarize
4. Tool result cache
4.1 What is it?
This is one of the most valuable caches in agentic systems. Suppose an agent calls:
weather_tool(city="Bangalore")and receives:
{
"temperature": 27,
"condition": "Cloudy"
}If another request asks for Bangalore's weather 30 seconds later, you may not need another external API call:
- Agent
- Tool
- Cache
- Hitreturn resultMisscall the external API
- Store result in cache
4.2 A generic tool cache
Instead of implementing caching separately for every tool, create a wrapper:
import hashlib
import json
def tool_cache_key(tool_name, arguments, tool_version):
payload = {
"tool": tool_name,
"version": tool_version,
"arguments": arguments,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"agent:tool:{digest}"Now:
def cached_tool_call(tool_name, arguments, tool_version, tool_function, ttl=300):
key = tool_cache_key(tool_name, arguments, tool_version)
cached = r.get(key)
if cached:
return json.loads(cached)
result = tool_function(**arguments)
r.setex(key, ttl, json.dumps(result))
return resultThis creates a reusable tool caching layer.
4.3 Deterministic vs non-deterministic tools
Not every tool should be cached the same way:
| Category | Examples |
|---|---|
| Highly cacheable | Currency conversion, historical weather, product specifications, company documentation, static metadata, public FAQ, geocoding |
| Short TTL | Current weather, stock prices, flight availability, hotel availability, inventory, exchange rates |
| Usually don't cache blindly | Create payment, place order, send email, delete account, transfer money, submit form |
The last group are side-effecting operations. Caching them can be dangerous.
4.4 Read tools vs write tools
A useful production classification:
| Tool type | Example | Treatment |
|---|---|---|
| Read | get_customer() |
Can potentially be cached |
| Write | create_customer() |
Needs idempotency and explicit execution semantics |
You should not run cache.get(key) before a write operation and assume that prevents duplicate execution. Instead, use an idempotency key + an execution record + transaction semantics.
5. API response cache
5.1 What is it?
Agents frequently call APIs: weather, search, maps, CRM. Every external API call adds latency, cost, rate-limit pressure and failure risk. A generic API cache can help:
def api_cache_key(service, endpoint, params, version="v1"):
payload = {
"service": service,
"endpoint": endpoint,
"params": params,
"version": version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"api:{digest}"Then:
def cached_api_request(service, endpoint, params, request_function, ttl=60):
key = api_cache_key(service, endpoint, params)
cached = r.get(key)
if cached:
return json.loads(cached)
response = request_function(endpoint, params)
r.setex(key, ttl, json.dumps(response))
return response6. Database query cache
6.1 What is it?
Agents may also query databases repeatedly (agent → SQL → database → result). A safe, read-only query can sometimes be cached:
def db_query_cache_key(database, query, parameters, schema_version):
payload = {
"database": database,
"query": query,
"parameters": parameters,
"schema": schema_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"db:query:{digest}"Then:
def cached_query(query, parameters, execute, ttl=30):
key = db_query_cache_key("production", query, parameters, "schema-v8")
cached = r.get(key)
if cached:
return json.loads(cached)
result = execute(query, parameters)
r.setex(key, ttl, json.dumps(result))
return resultBut database caching needs a careful invalidation strategy. After an UPDATE customer, the related query cache may be stale.
6.2 Cache database results by data version
A stronger approach is to keep a data version, for example customer_data_version = 481. When the relevant data changes, the version moves from 481 to 482, and because the version is part of the key (query + data version), the cache key changes with it.
This is often easier to reason about than trying to find every query affected by an update.
7. Function result cache
7.1 What is it?
Not every agent tool calls an external API. You may have expensive internal functions such as calculate_risk(), calculate_shipping(), generate_report(), classify_document() and score_customer(). These can be cached too:
def cached_function(function_name, function_version, arguments, function, ttl=600):
key = tool_cache_key(function_name, arguments, function_version)
cached = r.get(key)
if cached:
return json.loads(cached)
result = function(**arguments)
r.setex(key, ttl, json.dumps(result))
return resultThis turns caching into an infrastructure capability rather than an application-specific feature.
8. Agent trajectory cache
8.1 What is it?
An agent's execution can be represented as a trajectory:
- User
- Plan
- Tool A
- Tool B
- Tool C
- Final reasoning
We can store this trajectory. For example:
{
"goal": "Research product",
"steps": [
{ "tool": "search", "input": "product A", "output": "..." },
{ "tool": "pricing", "input": "product A", "output": "..." }
]
}This is useful for debugging, replay, evaluation, recovery, auditing and optimization.
But trajectory reuse must be handled carefully. A previous trajectory is not automatically valid in a new environment.
8.2 Reusing agent trajectories
Suppose an agent repeatedly receives:
"Generate our weekly sales report."
The workflow may be:
- Query the sales DB
- Query the returns DB
- Calculate metrics
- Generate a chart
- Summarize
Instead of letting the LLM rediscover this workflow every time, map the goal to the known workflow and run it on the current data:
- Goal
- Known workflow
- Execute on current data
This is much more reliable. Notice what we're caching: the workflow, not the old data. That's an important distinction.
9. Workflow step and partial execution cache
9.1 Workflow step cache
A better abstraction is to cache individual workflow steps. For the weekly report, if only step 1's input changes:
| Step | Action |
|---|---|
| 1. Fetch sales | Recompute |
| 2. Fetch returns | Reuse |
| 3. Calculate metrics | Reuse |
| 4. Generate chart | Reuse |
| 5. Summarize | Reuse |
This is much better than rerunning the entire agent.
9.2 Partial execution cache
This is one of the most powerful concepts in agentic caching. Suppose a task runs steps A → B → C → D → E, and execution fails at D.
- A
- B
- C
- D
- A ✓
- B ✓
- C ✓
- Resume from D
This can save enormous amounts of computation.
10. Agent checkpointing
10.1 Implementation
Store the execution state after important steps:
def checkpoint_key(execution_id, step_id):
return f"agent:checkpoint:{execution_id}:{step_id}"
def save_checkpoint(execution_id, step_id, state):
key = checkpoint_key(execution_id, step_id)
r.set(key, json.dumps(state))
def load_checkpoint(execution_id, step_id):
key = checkpoint_key(execution_id, step_id)
value = r.get(key)
if not value:
return None
return json.loads(value)Now a long-running agent can resume execution.
10.2 Checkpoint vs cache
These are related but not identical:
| Goal | Example | |
|---|---|---|
| Cache | Avoid recomputation | "Search result for query X" |
| Checkpoint | Recover execution state | "Agent execution 123 reached step 4 with these intermediate results" |
A production agent often needs both.
11. Agent context cache
11.1 What is it?
Agent context can become enormous: the system prompt, the conversation, the plan, tool outputs, RAG results and previous reasoning state. Instead of reconstructing everything, maintain structured context:
agent_context = {
"goal": "...",
"plan": [...],
"state": {...},
"tool_results": {...},
"retrieved_docs": [...],
"recent_events": [...],
}Then cache the context object:
def agent_context_key(session_id, context_version):
return f"agent:context:{session_id}:{context_version}"11.2 Context versioning
Suppose the agent context is at version 12. After a tool result arrives it becomes version 13, and after another tool call, version 14:
- Context v12
- Tool resultv13
- Another tool callv14
The cache key becomes agent:context:session123:v14. This is safer than mutating a single cache entry without tracking versions.
12. Agent memory cache
12.1 Short-term state and long-term memory
Long-running agents may have user preferences, past tasks, important facts, previous outcomes and learned patterns. Store durable memory separately from short-term execution state:
- Agent
- Short-term stateRedisLong-term memoryDB / vector store
Don't treat all memory as the same thing:
| Short-term | Long-term | |
|---|---|---|
| Holds | Current task, plan, tool results, conversation and execution state | User preferences, historical facts, past interactions, important persistent knowledge |
| Usually lives in | Redis | A database, vector store or other persistent storage |
Caching can speed up access to long-term memory without making Redis the source of truth.
13. Semantic tool cache
13.1 What is it?
Exact tool caching handles search("latest AI news"). A semantic tool cache might also treat "recent AI news" as the same request.
However, this should only be used when semantic equivalence is safe. For a search tool, "What happened in AI today?" and "What happened in AI yesterday?" are semantically close but need different results.
So semantic tool caching needs semantic similarity + time constraints + tool semantics + freshness.
14. Cache-aware tools
14.1 Cache-aware tool design
A powerful production pattern is to make tools explicitly cache-aware. Instead of exposing only search(query), define search(query, cache_policy), for example:
cache_policy = {
"enabled": True,
"ttl": 300,
"scope": "tenant",
"freshness": "5m",
}The agent runtime can then apply consistent caching rules.
14.2 Tool metadata should describe cacheability
A tool registry could contain:
tools = {
"weather": {
"cacheable": True,
"ttl": 300,
"side_effect": False,
},
"get_customer": {
"cacheable": True,
"ttl": 30,
"side_effect": False,
},
"send_email": {
"cacheable": False,
"side_effect": True,
},
}Now the agent runtime knows how to treat each tool. This is much safer than letting every agent developer invent their own rules.
15. Side effects: idempotency and retries
15.1 Idempotency for write tools
Caching isn't enough for side-effecting operations. Suppose the agent calls send_email() and the request times out. The agent doesn't know whether the email was sent, so it retries, and two emails may go out.
The right solution is often idempotency, not caching:
idempotency_key = f"{execution_id}:{tool_name}:{step_id}"The server stores the idempotency key with the execution result. Retries then return the original result instead of executing the operation again.
15.2 Retry and cache interaction
When a tool call hits an API timeout, before retrying the system should determine whether the operation was:
- read-only
- partially completed
- side-effecting
- idempotent
For reads, a plain retry is usually straightforward. For writes, retry + idempotency is much safer.
16. Multi-agent cache sharing
16.1 Sharing work between agents
Now imagine a supervisor with three agents that all need the same information:
- Supervisor agent
- Research agentData agentWriter agent
- Shared tool cache
Without shared caching, each agent calls the API itself. With a shared tool cache, the first call is reused by the others. This can significantly reduce duplicated work.
But shared caches need strict scope.
16.2 What can be shared across agents?
| Potentially shareable | Usually not blindly shareable |
|---|---|
| Public search results | User-specific data |
| Public documentation | Private customer records |
| Product metadata | Authorization-sensitive results |
| Static configuration | Personalized recommendations |
| Common embeddings | Private tool outputs |
| Common reference data |
The cache scope should be explicit: global, tenant, team, user, session or execution.
17. Agent response cache
17.1 What is it?
At the top level, we can still cache final responses. But the response cache key should include the agent's important dependencies:
def agent_response_key(
tenant_id,
user_scope,
task,
agent_version,
toolset_version,
context_version,
):
payload = {
"tenant": tenant_id,
"scope": user_scope,
"task": task,
"agent": agent_version,
"tools": toolset_version,
"context": context_version,
}
digest = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
return f"agent:response:{digest}"17.2 Why agent response caching is harder
An agent response can depend on the user, the time, tools, external APIs, the database, memory, RAG, the agent version, tool versions and execution state. So for agents especially:
Same prompt ≠ same answer.
"What's the weather?" should almost never have a long-lived response cache. But "What are the steps in our internal onboarding process?" may be highly cacheable.
18. Dependencies, versions and invalidation
18.1 The agent cache dependency graph
A production agent can have a dependency graph like:
- User request
- Agent plan
- Tool ATool BTool C
- Result AResult BResult C
- Agent context
- Next step
- Final answer
If Tool B's result changes, everything downstream of it may need recomputing:
- Tool B
- Result B
- Context
- Next step
- Final answer
This is why dependency tracking becomes more important as agents become more complex.
18.2 Cache invalidation in agents
There is no universal TTL that solves agent caching. Instead, ask:
What caused this result to become invalid?
| Cause | Version change |
|---|---|
| Tool version changed | tool-v1 → tool-v2 |
| Agent prompt changed | prompt-v4 → prompt-v5 |
| External data changed | data-v20 → data-v21 |
| User state changed | cart-v8 → cart-v9 |
| Permissions changed | role-v2 → role-v3 |
These versions can be part of cache keys.
18.3 Version everything important
A useful production pattern is to version the agent, prompts, tools, workflows, knowledge, schema and data. Then:
cache key = input + agent version + prompt version + tool version + data version
Changing a version naturally creates a new cache namespace.
19. Long-running agents and recovery
19.1 Long-running agents
Some agents don't finish in seconds. Research agents, data analysis agents, report generators, software engineering agents and workflow automation agents may run for minutes or hours.
For them, caching is less about response reuse and more about execution recovery. This is where checkpoints become critical.
19.2 Production agent execution
Instead of:
result = run_agent(task)think:
execution = start_execution(task)
while not execution.complete:
step = execution.next_step()
result = execute_step(step)
save_checkpoint(execution.id, step.id, execution.state)Now if the process crashes, it restarts, loads the checkpoint and resumes, rather than starting from zero.
19.3 Failure recovery
Imagine steps 1–3 succeeded, step 4 failed and step 5 never started. Recovery should look like:
- Load checkpoint
- Retry step 4
- Step 5
- Final response
not a rerun of steps 1 through 5. This is one of the biggest benefits of caching intermediate execution state.
20. Agent caching meets RAG caching
20.1 Agentic RAG
Agentic RAG combines both systems:
- Agent
- Query transformation cache
- Embedding cache
- Retrieval cache
- Reranking cache
- Agent tool cache
- Agent state
- Agent context
- LLM
Now caching happens across several dimensions:
| Area | Caches |
|---|---|
| RAG | Embedding cache, retrieval cache, reranking cache |
| Tools | API cache, DB cache, function cache |
| State | Session state, execution state |
| Memory | Short-term, long-term |
This is where a unified cache architecture becomes valuable.
21. A shared agent cache layer
21.1 A generic agent cache wrapper
We can create one reusable abstraction:
import json
class AgentCache:
def __init__(self, redis_client):
self.redis = redis_client
def get(self, key):
value = self.redis.get(key)
if not value:
return None
return json.loads(value)
def set(self, key, value, ttl):
self.redis.setex(key, ttl, json.dumps(value))
def get_or_compute(self, key, compute, ttl=300):
cached = self.get(key)
if cached is not None:
return cached
result = compute()
self.set(key, result, ttl)
return resultUsage:
cache = AgentCache(r)
result = cache.get_or_compute(
key="agent:tool:weather:bangalore",
compute=lambda: weather_tool("Bangalore"),
ttl=300,
)This creates a consistent caching interface across the agent runtime.
21.2 Add a cache policy
A more production-ready version accepts a policy:
cache_policy = {
"enabled": True,
"ttl": 300,
"scope": "tenant",
"stale_while_revalidate": True,
"semantic": False,
}The runtime can then decide:
- Should I cache?
- For how long?
- Who can reuse it?
- Can stale data be returned?
- Can semantic matching be used?
This moves caching from ad-hoc code into an explicit system design.
22. Observability
22.1 Observability for agent caching
Traditional cache metrics aren't enough. Track agent.cache.hit, agent.cache.miss, agent.cache.stale and agent.cache.lock_wait, but also:
agent.tool.calls_avoidedagent.llm.calls_avoidedagent.steps_reusedagent.steps_recomputedagent.tokens_savedagent.execution_time_savedagent.cost_saved
For example:
metrics.increment("agent.tool.cache_hit", tags={"tool": tool_name})
metrics.increment("agent.execution.step_reused")22.2 Measure agent-level savings
Suppose a typical agent run looks like this before and after caching:
| Metric | Before caching | After caching |
|---|---|---|
| Tool calls | 12 | 4 |
| LLM calls | 8 | 3 |
| Latency | 25 s | 9 s |
| Cost | $0.40 | $0.12 |
The important metric isn't just "cache hit rate = 70%". The real business metrics are cost, from $0.40 to $0.12, and latency, from 25 s to 9 s. That's what matters.
22.3 Don't optimize for cache hit rate alone
A 99% hit rate can still be bad. A 99% hit rate on a cheap operation may be worth far less than a 10% hit rate on an expensive LLM operation.
A better metric:
Cache value = avoided computation cost × reuse frequency × correctness confidence
23. Security and scope
23.1 Security considerations
Agent caching introduces significant security risks. Never accidentally share:
- user secrets
- API responses
- private documents
- authorization-sensitive results
- customer records
- tool credentials
- private agent memory
And never cache credentials (API keys, passwords, OAuth tokens, session tokens) as if they were normal tool results.
23.2 Cache scope
Every cache should answer:
Who is allowed to reuse this result?
The possible scopes are global, tenant, team, user, session and execution. For example:
| Data | Scope |
|---|---|
| Public documentation | Global |
| Company pricing | Tenant |
| User preferences | User |
| Current workflow state | Execution |
This makes cache boundaries explicit.
24. Before you ship
24.1 Production agent caching checklist
State
- Agent state is explicit
- Session state is cached
- Execution state is versioned
- Checkpoints exist for long-running tasks
Planning
- Planning cache evaluated
- Agent version included
- Prompt version included
- Dynamic plans aren't blindly reused
Tools
- Tools classified as read or write
- Cacheability defined
- TTL defined per tool
- Tool version included
APIs
- API response caching evaluated
- Freshness requirements defined
- Rate limits considered
- Errors aren't cached blindly
Databases
- Read queries cached where appropriate
- Data versioning considered
- Mutation invalidation defined
Execution
- Workflow steps are independently cacheable
- Partial results can be reused
- Checkpointing implemented
- Failed executions can resume
Multi-agent
- Shared cache boundaries defined
- Tenant isolation enforced
- User-specific results protected
Security
- Authorization scope included
- Sensitive data protected
- Credentials never cached as ordinary results
- Cache access monitored
Observability
- Hit rate
- Miss rate
- Steps reused
- Tool calls avoided
- LLM calls avoided
- Tokens saved
- Cost saved
- Latency saved
25. The full picture
25.1 The complete agentic cache architecture
Putting everything together:
- User
- Response cacheon a miss, continue
- Agent state cache
- Planning cache
- Agent loop
- Tool cacheRAG cacheAPI cache
- Function cacheRetrieval cacheExternal API
- Intermediate results
- Agent context cache
- Checkpoint storecontinue the loop until done
- Final answer
25.2 The big shift in thinking
Traditional caching asks:
"Can I cache the response?"
RAG caching asks:
"Can I cache the retrieval?"
Agentic caching asks a much bigger question:
"Which parts of the agent's work can I safely reuse?"
That could be the plan, a tool call, an API result, a database result, a search result, an embedding, a reranking, an intermediate calculation, a workflow step, agent state, agent memory, context, a checkpoint or the final response.
The agent doesn't always need to start from zero.
25.3 The production principle
A naive agent does everything on every request:
- Every request
- Think
- Plan
- Call tools
- Retrieve
- Compute
- Think again
- Answer
A production agent does this instead:
- Every request
- What have I already computed?
- What is still valid?
- What has changed?
- Reuse valid work
- Compute only the delta
- Continue execution
That's the real value of agentic caching.
Don't cache just the answer. Cache the work that produced the answer.
25.4 The final mental model
Think of an agent execution as a computation graph:
- User request
- Plan
- SearchAPIDB
- Result AResult BResult C
- Analysis
- Decision
- Tool
- Answer
Each node is potentially a cacheable computation:
| If… | Then… |
|---|---|
| The search has already run and is still valid | Reuse it |
| The API data has changed | Recompute it |
| The DB result is user-specific | Scope it correctly |
| The tool has side effects | Use idempotency |
| The execution fails | Resume from the checkpoint |
That is how caching evolves from a simple optimization into an execution architecture.
25.5 Part 5 core takeaways
If you remember only ten things:
- Cache agent state.
- Cache deterministic tool results.
- Separate read tools from write tools.
- Use idempotency for side effects.
- Cache workflow steps, not only final responses.
- Checkpoint long-running executions.
- Version agents, prompts, tools and data.
- Make cache scope explicit.
- Track work avoided, not just cache hits.
- Design agents to reuse previous computation.
The ultimate goal is simple:
A production agent should not redo work that is still valid.
And that leads to the next level of the caching architecture.
What's next?
Part 6: LLM inference caching
We've now covered caching around the application: conversation, RAG, agents and tools. But there is another layer where enormous amounts of computation happen, inside the model:
- LLM
- PrefillDecode
- KV cacheTokens
Part 6 will go deeper into:
- KV cache, prefix cache, prompt caching and prefill caching
- decode-time caching, attention KV reuse and session-level KV reuse
- prefix matching, cache-friendly prompt design, and static vs dynamic prompt sections
- long-context caching, multi-turn KV reuse and batch inference caching
- speculative decoding and continuous batching
- GPU memory constraints, KV cache quantization and KV cache eviction
- paged attention and prefix sharing
- inference cost optimization, latency vs memory trade-offs, and production LLM serving architecture
The central question becomes:
What if we could avoid making the model recompute the same attention work every time?
That takes us from application-level caching into the internals of LLM inference itself.
Next: Part 6, LLM inference caching covers the KV cache, prefix caching and prompt caching inside the model.