Building an Agent from First Principles
Build an AI agent from scratch in Python: tools, the observe-decide-act loop, state, planning, error handling, termination and verification.
Topic
Working with large language models in real software, from API calls to production.
Build an AI agent from scratch in Python: tools, the observe-decide-act loop, state, planning, error handling, termination and verification.
Beyond better prompts: how to design, assemble, optimize and manage the information an LLM needs to solve a task.
A deep technical guide to LLM internals, from raw text to token generation, with Python implementations from scratch.
Bringing it all together: how to design caching as a first-class part of an AI system, from keys and versions to reliability, cost and rollout.
Measure what the cache actually saves: latency, tokens, calls avoided, cost, ROI, hit quality and freshness, not just the hit rate.
Fast is useless if the answer is wrong: TTLs, versioned keys, dependency graphs, events, concurrency and testing for correct AI caches.
Scaling AI caches: Redis Cluster sharding, hot keys, L1/L2 caches, stampedes, invalidation at scale, multi-region, failure handling and cost.
KV cache, prefix cache and prompt caching inside LLM inference: how they differ, what they cost in GPU memory, and how to design prompts for reuse.
Agent state, plans, tool and API results, workflow steps, checkpoints, idempotent writes and scoped sharing: caching the work an agent does.
Make retrieval faster, cheaper and smarter: caching queries, embeddings, retrieval, reranking, chunks, context and answers without serving stale data.
Caching for chat: conversation history, sessions, summaries, user memory, context fingerprints and responses, without leaking or going stale.
Caching for agents: tool results, workflow steps, state, checkpoints, MCP tools and resources, idempotency, and keeping it all safe.
KV, prompt, semantic, embedding, retrieval, reranking, tool, API, conversation and agent state caches: what each saves and how to key it safely.
From demo to dependable: retrieval, reranking, grounding, citations and evaluation for retrieval-augmented generation in production.