Building an Agent from First Principles
Build an AI agent from scratch in Python: tools, the observe-decide-act loop, state, planning, error handling, termination and verification.
Blog
14 posts, newest first
Build an AI agent from scratch in Python: tools, the observe-decide-act loop, state, planning, error handling, termination and verification.
Beyond better prompts: how to design, assemble, optimize and manage the information an LLM needs to solve a task.
A deep technical guide to LLM internals, from raw text to token generation, with Python implementations from scratch.
Bringing it all together: how to design caching as a first-class part of an AI system, from keys and versions to reliability, cost and rollout.
Measure what the cache actually saves: latency, tokens, calls avoided, cost, ROI, hit quality and freshness, not just the hit rate.
Fast is useless if the answer is wrong: TTLs, versioned keys, dependency graphs, events, concurrency and testing for correct AI caches.
Scaling AI caches: Redis Cluster sharding, hot keys, L1/L2 caches, stampedes, invalidation at scale, multi-region, failure handling and cost.
KV cache, prefix cache and prompt caching inside LLM inference: how they differ, what they cost in GPU memory, and how to design prompts for reuse.
Agent state, plans, tool and API results, workflow steps, checkpoints, idempotent writes and scoped sharing: caching the work an agent does.
Make retrieval faster, cheaper and smarter: caching queries, embeddings, retrieval, reranking, chunks, context and answers without serving stale data.