Context Engineering Is the New Prompt Engineering
Beyond better prompts: how to design, assemble, optimize and manage the information an LLM needs to solve a task.
Topic
Retrieval-augmented generation, search and grounding answers in your own data.
Beyond better prompts: how to design, assemble, optimize and manage the information an LLM needs to solve a task.
Bringing it all together: how to design caching as a first-class part of an AI system, from keys and versions to reliability, cost and rollout.
Fast is useless if the answer is wrong: TTLs, versioned keys, dependency graphs, events, concurrency and testing for correct AI caches.
Make retrieval faster, cheaper and smarter: caching queries, embeddings, retrieval, reranking, chunks, context and answers without serving stale data.
KV, prompt, semantic, embedding, retrieval, reranking, tool, API, conversation and agent state caches: what each saves and how to key it safely.
From demo to dependable: retrieval, reranking, grounding, citations and evaluation for retrieval-augmented generation in production.