HaribaskarAI Engineer
← All posts

Building an Agent from First Principles

Build an AI agent from scratch in Python: tools, the observe-decide-act loop, state, planning, error handling, termination and verification.

Haribaskar Dhanabalan27 min read

A robot labelled "AI Agent" working at a desk with a holographic AI dashboard, in front of a rainy neon city
On this page

AI Engineering series, part 3. Earlier parts: What Actually Happens Between Your Prompt and the Next Token? and Context Engineering Is the New Prompt Engineering.

An LLM can answer questions, summarize documents, generate code, and explain complex concepts. But answering a question is fundamentally different from completing a task that requires multiple decisions, external actions, and verification.

Consider this request:

Investigate why the payment API is failing, identify the root cause, implement a fix, run the relevant tests, and report the result.

A normal LLM call can suggest debugging steps or generate a patch. But completing the task requires much more:

  • Inspecting the environment.
  • Deciding which action to take next.
  • Calling tools.
  • Observing their results.
  • Updating its plan based on new evidence.
  • Managing execution state.
  • Recovering from failures.
  • Knowing when to stop.

This is where AI agents become useful.

The interesting part is that an agent does not necessarily require a complex multi-agent framework, a specialized model, or an elaborate architecture. At its core, an agent can be built from a model, a set of tools, a state representation, and an execution loop.

In this article, we'll build an agent from first principles, starting with the theory and progressing to a practical Python implementation. We'll implement tool calling, observations, state management, execution limits, error handling, and a basic planning loop. Finally, we'll examine what separates a working prototype from a reliable production agent.

  1. Observe → Decide → Act → Observe again
A useful mental model

An agent repeatedly observes the environment, chooses an action, executes it, and uses the result to decide what happens next.

1. What exactly is an AI agent?

An AI agent is a system that uses observations of its environment to choose and execute actions in pursuit of a goal.

For an LLM-based agent, the language model typically helps interpret the task, select actions, or decide what to do next. The surrounding software provides the tools, state, execution control, and safety boundaries.

A useful abstraction is:

a_t = π(o_t, s_t, g)

Where:

  • a_t is the action selected at step t.
  • o_t is the current observation.
  • s_t is the agent's state.
  • g is the overall goal.
  • π is the agent's decision policy.

After the action executes, the environment produces a new observation:

o_(t+1) = E(a_t)

Here, E represents the environment's response to the action. In practice, the outcome may also depend on external state, timing, permissions, and failures.

The agent then evaluates the new observation and decides whether to continue, change direction, or terminate.

This creates a feedback loop.

1.1 An agent versus a normal LLM call

Capability Basic LLM call LLM-based agent
Generates text Yes Yes
Selects actions dynamically Usually not by itself Yes
Uses external tools Only if the application orchestrates them Core capability
Observes execution results Not automatically Feeds results back into decisions
Maintains task state Application-dependent Explicitly managed by the agent runtime
Adapts after failures Requires additional application logic Can use observations to choose recovery actions
Controls when to stop Often one response or a fixed loop Uses explicit termination rules

An agent is not necessarily more intelligent than the underlying LLM. Its advantage comes from combining model-driven decisions with real-world information, actions, and feedback.

The key distinction is the feedback loop, not the presence of a prompt that says "You are an agent."

2. The five fundamental components of an agent

A useful first-principles architecture consists of five components.

Component Responsibility Example
1. Goal Defines the desired outcome and constraints. Identify the cause of a failing test without modifying unrelated files.
2. Decision policy Uses the LLM, rules, or another policy to select the next action. Inspect a log before attempting a code change.
3. Tools Allow the agent to interact with its environment. Read a file, run a test, search documentation, or query a database.
4. State and observations Preserve task progress, tool results, decisions, and outstanding work. The failing test has been identified, but the integration test remains unverified.
5. Execution controller Enforces tool permissions, step limits, timeouts, validation, and termination rules. Stop after a bounded number of iterations or when a tool is not authorized.

These components do not need to be separate services. For a small agent, they can be ordinary Python functions and data structures.

The important thing is that their responsibilities are clearly defined.

3. The agent loop: Observe → Decide → Act → Observe

The fundamental agent loop can be represented as:

  1. Receive goal
  2. Observe current state
  3. LLM selects next action
  4. Action valid?no → reject or request a correction, then decide again
  5. Execute authorized tool
  6. Capture result
  7. Update state
  8. Goal satisfied?no → observe again and repeat
  9. Return verified result
The agent loop

Let's examine each stage.

3.1 Observe

The agent collects the information required to make its next decision.

Observations can include:

  • The user's request.
  • Tool results.
  • Files and application state.
  • Retrieved documents.
  • Test results.
  • Previous decisions.
  • Errors and warnings.

An observation should represent what the system actually knows, not what it assumes happened.

3.2 Decide

The decision policy determines the next action. For example:

Observation:
The unit test fails because permanent validation errors
are being retried.
 
Possible actions:
1. Inspect the retry handler.
2. Change the retry limit.
3. Disable retries.
4. Report the failure.

A capable agent should select an action based on the available evidence and task constraints.

The LLM can help make this choice, but deterministic rules can also be used for simple or safety-critical decisions.

3.3 Act

The selected tool is executed by the application. Examples:

  • Reading a file.
  • Running a test.
  • Searching a knowledge base.
  • Calling an API.
  • Creating a ticket.

The model proposes an action; the runtime decides whether that action is valid and authorized before executing it.

3.4 Observe again

The tool result becomes new evidence.

If the test passes, the agent may proceed to broader verification. If the test fails, it may inspect the error and revise its approach.

This feedback loop is what makes the agent adaptive.

4. Build the simplest possible agent without an LLM

Before introducing an LLM, let's implement a deterministic agent. This helps separate the agent architecture from the model itself.

Imagine a robot in a small grid world. It must reach a target while avoiding obstacles.

from dataclasses import dataclass, field
 
Position = tuple[int, int]
 
 
@dataclass
class GridEnvironment:
    position: Position
    goal: Position
    obstacles: set[Position] = field(default_factory=set)
 
    def observe(self) -> dict:
        return {
            "position": self.position,
            "goal": self.goal,
            "obstacles": self.obstacles,
        }
 
    def step(self, action: str) -> dict:
        x, y = self.position
        moves = {
            "up": (x, y - 1),
            "down": (x, y + 1),
            "left": (x - 1, y),
            "right": (x + 1, y),
        }
 
        if action not in moves:
            return {"ok": False, "error": "Invalid action"}
 
        next_position = moves[action]
        if next_position in self.obstacles:
            return {"ok": False, "error": "Obstacle detected"}
 
        self.position = next_position
        return {"ok": True, "position": self.position}

The environment holds the state and applies actions. The agent needs a policy that turns an observation into an action, and a loop that repeats until a stopping condition is reached:

def greedy_policy(observation: dict) -> str:
    """Move horizontally towards the goal, then vertically."""
    (x, y), (goal_x, goal_y) = observation["position"], observation["goal"]
    if x < goal_x:
        return "right"
    if x > goal_x:
        return "left"
    if y < goal_y:
        return "down"
    return "up"
 
 
def run_grid_agent(env: GridEnvironment, max_steps: int = 20) -> dict:
    for step in range(max_steps):
        observation = env.observe()
        if observation["position"] == observation["goal"]:
            return {"status": "completed", "steps": step}
 
        action = greedy_policy(observation)
        result = env.step(action)
        if not result["ok"]:
            return {"status": "failed", "steps": step + 1, "reason": result["error"]}
 
    return {"status": "max_steps_reached", "steps": max_steps}
 
 
print(run_grid_agent(GridEnvironment(position=(0, 0), goal=(3, 2))))
# {'status': 'completed', 'steps': 5}
 
print(run_grid_agent(GridEnvironment(position=(0, 0), goal=(3, 2), obstacles={(2, 0)})))
# {'status': 'failed', 'steps': 2, 'reason': 'Obstacle detected'}

This example illustrates several fundamental ideas:

  • The environment has a state.
  • The agent receives observations.
  • A policy chooses actions.
  • Actions change the environment.
  • The agent repeats until a stopping condition is reached.

The policy is deterministic and deliberately simple. It is not a general pathfinding algorithm, and as the second run shows, it gets stuck because it does not plan around obstacles.

That limitation is useful: it demonstrates why an agent needs a policy capable of handling the actual complexity of its environment.

For our next implementation, the policy will be an LLM that selects tools based on observations.

5. Introduce tools: give the LLM the ability to act

A model that can only generate text cannot independently inspect a local file or execute a test. The application must expose those capabilities through tools.

A tool generally has three parts:

  1. A name.
  2. An input schema.
  3. An implementation.

For example, a coding agent might have:

read_file(path)
    → Returns file contents
 
search_files(query)
    → Returns matching files
 
run_tests(test_path)
    → Returns test output and exit status

The model decides which tool to request and what arguments to provide. The runtime validates the request, executes the tool, and returns its result.

5.1 Implement a small tool registry

Let's create a reusable registry.

from dataclasses import dataclass
from typing import Any, Callable
 
 
@dataclass
class Tool:
    name: str
    description: str
    function: Callable[..., Any]
 
 
class ToolRegistry:
    def __init__(self):
        self._tools: dict[str, Tool] = {}
 
    def register(self, tool: Tool):
        if tool.name in self._tools:
            raise ValueError(f"Tool already registered: {tool.name}")
        self._tools[tool.name] = tool
 
    def get(self, name: str) -> Tool:
        if name not in self._tools:
            raise ValueError(f"Unknown tool: {name}")
        return self._tools[name]
 
    def describe(self) -> list[dict]:
        return [
            {"name": tool.name, "description": tool.description}
            for tool in self._tools.values()
        ]

This registry provides a central place to discover and execute tools.

It deliberately does not yet handle schemas, authorization, timeouts, or structured errors. We will add validation in later sections.

5.2 Register example tools

For a first demonstration, we'll use a simulated environment rather than giving the agent unrestricted access to the host machine.

files = {
    "payment/retry.py": (
        "def should_retry(error):\n"
        "    return True\n"
    ),
    "payment/tests/test_retry.py": (
        "def test_validation_errors_are_not_retried():\n"
        '    assert should_retry("validation_error") is False\n'
    ),
}
 
 
def read_file(path: str) -> dict:
    if path not in files:
        return {"ok": False, "error": "File not found"}
    return {"ok": True, "path": path, "content": files[path]}
 
 
def list_files() -> dict:
    return {"ok": True, "files": list(files.keys())}
 
 
registry = ToolRegistry()
registry.register(
    Tool(
        name="read_file",
        description="Read a project file using its exact project-relative path.",
        function=read_file,
    )
)
registry.register(
    Tool(
        name="list_files",
        description="List available project files.",
        function=list_files,
    )
)

Notice that read_file reads only from a predefined dictionary. It cannot access arbitrary filesystem paths.

This is intentional. An agent's tool surface should be as narrow as practical.

For real applications, filesystem access, shell commands, network access, and database writes require their own permission checks.

6. Implement the decision–execution–observation loop

We now have an environment and a tool registry. Next, let's define how an agent action is represented.

A basic action might look like this:

action = {
    "type": "tool_call",
    "tool": "read_file",
    "arguments": {
        "path": "payment/retry.py",
    },
}

The runtime executes the action and returns an observation.

def execute_action(registry: ToolRegistry, action: dict) -> dict:
    if action.get("type") != "tool_call":
        return {"ok": False, "error": "Unsupported action type"}
 
    tool_name = action.get("tool")
    arguments = action.get("arguments", {})
 
    try:
        tool = registry.get(tool_name)
        if not isinstance(arguments, dict):
            return {"ok": False, "tool": tool_name, "error": "Arguments must be an object"}
 
        result = tool.function(**arguments)
        return {"ok": True, "tool": tool_name, "result": result}
    except (ValueError, TypeError) as exc:
        return {"ok": False, "tool": tool_name, "error": str(exc)}

Test it:

observation = execute_action(registry, action)
print(observation)

Expected result:

{
    'ok': True,
    'tool': 'read_file',
    'result': {
        'ok': True,
        'path': 'payment/retry.py',
        'content': 'def should_retry(error):\n    return True\n',
    },
}

The architecture now has a clear boundary:

  • The policy chooses an action.
  • The runtime executes it.
  • The result becomes an observation.
  • The policy receives that observation on the next iteration.

In production, do not catch every exception and return an unclassified string. Use structured error types, record diagnostic details securely, and distinguish retryable failures from invalid requests and permanent errors.

7. Add an LLM: let the model choose the next action

Now we can introduce the model.

We will use the OpenAI Python SDK's tool-calling interface. Tool calling allows the model to request a function invocation using structured arguments.

The model does not execute the Python function itself. Your application receives the tool request and decides what to do with it.

Install the SDK if necessary:

pip install openai

Configure your credentials through environment variables:

export OPENAI_API_KEY="your-api-key"
export OPENAI_MODEL="your-available-model"

Use a model that supports tool calling in your account.

7.1 Define the tool schemas

The API needs descriptions of the available tools and their accepted arguments.

TOOLS = [
    {
        "type": "function",
        "name": "list_files",
        "description": "List available project files.",
        "parameters": {
            "type": "object",
            "properties": {},
            "required": [],
            "additionalProperties": False,
        },
        "strict": True,
    },
    {
        "type": "function",
        "name": "read_file",
        "description": "Read a file using its exact project-relative path.",
        "parameters": {
            "type": "object",
            "properties": {
                "path": {
                    "type": "string",
                    "description": "Project-relative file path",
                },
            },
            "required": ["path"],
            "additionalProperties": False,
        },
        "strict": True,
    },
]

A schema helps constrain the structure of the arguments, but it does not establish authorization or guarantee that a tool call is safe.

The application still needs to validate the requested operation.

7.2 Build the agent loop

The following implementation uses the Responses API and repeatedly processes tool calls until the model returns a final response or the step limit is reached.

import json
import os
 
from openai import OpenAI
 
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
MODEL = os.environ["OPENAI_MODEL"]
 
INSTRUCTIONS = (
    "You are a debugging assistant. Use the tools to inspect the project. "
    "Base every conclusion on tool results, and say what you could not verify."
)
 
 
def dispatch_tool(name: str, arguments: dict) -> dict:
    try:
        tool = registry.get(name)
        if not isinstance(arguments, dict):
            return {"ok": False, "error": "Arguments must be an object"}
        result = tool.function(**arguments)
        return {"ok": True, "result": result}
    except (ValueError, TypeError) as exc:
        return {"ok": False, "error": str(exc)}
 
 
def run_agent(goal: str, max_steps: int = 8) -> str:
    conversation: list = [{"role": "user", "content": goal}]
 
    for _ in range(max_steps):
        response = client.responses.create(
            model=MODEL,
            instructions=INSTRUCTIONS,
            input=conversation,
            tools=TOOLS,
        )
 
        tool_calls = [item for item in response.output if item.type == "function_call"]
        if not tool_calls:
            return response.output_text  # the model produced its final answer
 
        # Keep the model's tool requests in the conversation, then add each result.
        conversation += response.output
        for call in tool_calls:
            try:
                arguments = json.loads(call.arguments)
            except json.JSONDecodeError:
                result = {"ok": False, "error": "Arguments were not valid JSON"}
            else:
                result = dispatch_tool(call.name, arguments)
 
            conversation.append(
                {
                    "type": "function_call_output",
                    "call_id": call.call_id,
                    "output": json.dumps(result),
                }
            )
 
    return f"Stopped: the step limit of {max_steps} was reached before a final answer."
 
 
print(run_agent(
    "Find out why test_validation_errors_are_not_retried fails. "
    "Inspect the relevant files and explain the root cause."
))

This is a small tool-using agent, not a fully autonomous software engineer. Its tools can inspect the simulated files but cannot modify them or execute tests.

The loop demonstrates the essential execution pattern:

  1. Send the goal and available tool schemas.
  2. Receive a model response.
  3. Detect requested tool calls.
  4. Validate and execute the requested operations.
  5. Return their outputs to the model.
  6. Continue until the model produces a response or the step limit is reached.

The exact SDK interfaces and model capabilities can vary by version. Consult the current API documentation when adapting this example to a specific deployment.

8. What happens internally during a tool call?

Understanding the protocol is important because tool calling is one of the most misunderstood parts of agent development.

Imagine the user asks:

Find the failing retry logic and inspect its implementation.

The model may request:

{
  "name": "read_file",
  "arguments": {
    "path": "payment/retry.py"
  }
}

The application then performs the operation and returns:

{
  "ok": true,
  "result": {
    "path": "payment/retry.py",
    "content": "def should_retry(error): return True"
  }
}

The model receives the result and may request another tool or produce its final answer.

  1. User → Agent runtime"Investigate retry logic"
  2. Agent runtime → LLMgoal + tool definitions
  3. LLM → Agent runtimeread_file(path)
  4. Agent runtimevalidates the tool and its arguments
  5. Agent runtime → Toolexecute read_file
  6. Tool → Agent runtimefile contents
  7. Agent runtime → LLMtool result
  8. LLM → Agent runtimea diagnosis, or another tool call
  9. *Agent runtime → User*final response
One tool call, end to end

The key architectural distinction is that the LLM selects the action, but the application owns execution.

This separation enables:

  • Permission enforcement.
  • Input validation.
  • Tool-call logging.
  • Timeout handling.
  • Retries for appropriate failures.
  • Human approval for sensitive operations.
  • Deterministic testing of the execution layer.

It also allows the same agent policy to operate against different tool implementations, provided the contracts remain compatible.

9. Agent state: why a conversation transcript is not enough

So far, our agent loop has been mostly stateless apart from the conversation maintained through the API.

For simple tasks, that can be sufficient. For longer workflows, the application should maintain explicit task state.

Consider an agent tasked with fixing a failing test. It might need to remember:

  • Which files it inspected.
  • What it believes caused the failure.
  • Which changes it attempted.
  • Which tests passed.
  • Which tests remain unverified.
  • What action should happen next.

Representing these facts as structured state makes the agent easier to recover, evaluate, and debug.

9.1 Define a task state

from dataclasses import dataclass, field
 
 
@dataclass
class AgentState:
    goal: str
    current_step: int = 0
    observations: list[dict] = field(default_factory=list)
    completed_actions: list[dict] = field(default_factory=list)
    pending_actions: list[str] = field(default_factory=list)
    status: str = "running"
    final_answer: str | None = None

Now update the state after each action.

def record_observation(state: AgentState, observation: dict):
    state.observations.append(observation)
 
 
def record_action(state: AgentState, action: dict):
    state.completed_actions.append(action)
    state.current_step += 1
 
 
def finish_agent(state: AgentState, answer: str):
    state.status = "completed"
    state.final_answer = answer

These functions are simple, but they establish a useful distinction:

  • Conversation history records messages exchanged with the model.
  • Execution state records what the application knows about the task.
  • External artifacts preserve files, logs, reports, and other evidence.

These should not be treated as interchangeable.

For example, a model-generated summary might say that a test passed. The execution state should record that claim as verified only when a test runner actually returned a successful result.

9.2 Build a context from state

Instead of sending every previous observation, construct the next model input from the current goal and the relevant state.

def state_to_context(state: AgentState) -> str:
    return f"""
Goal:
{state.goal}
 
Completed actions:
{state.completed_actions}
 
Recent observations:
{state.observations[-5:]}
 
Pending actions:
{state.pending_actions}
 
Current status:
{state.status}
""".strip()

This is a basic example of context engineering applied to agents.

In production, observations should have explicit size limits and provenance. Large tool results should usually be stored separately and referenced by identifiers rather than inserted wholesale into every model call.

10. Planning: should the agent create a plan before acting?

A basic agent can choose the next action directly from the current observation. This is often called reactive behavior.

For a complex task, the agent may benefit from decomposing the goal into smaller steps. For example:

Goal:
Investigate and fix the payment retry failure.
 
Plan:
1. Inspect the failing test.
2. Read the retry handler.
3. Identify the incorrect behavior.
4. Propose a minimal code change.
5. Apply the change.
6. Run the targeted test.
7. Verify the final result.

A plan helps the system organize the task, but it should not be treated as an immutable script. New evidence may invalidate an earlier assumption.

10.1 A simple plan representation

@dataclass
class PlanStep:
    description: str
    status: str = "pending"
    evidence: list[str] = field(default_factory=list)
 
 
@dataclass
class AgentPlan:
    goal: str
    steps: list[PlanStep]

Create a plan:

plan = AgentPlan(
    goal="Fix the payment retry failure",
    steps=[
        PlanStep("Inspect the failing test"),
        PlanStep("Inspect the retry handler"),
        PlanStep("Identify the root cause"),
        PlanStep("Implement the minimal fix"),
        PlanStep("Run targeted tests"),
        PlanStep("Report verified results"),
    ],
)

A controller can select the next pending step:

def next_pending_step(plan: AgentPlan) -> PlanStep | None:
    for step in plan.steps:
        if step.status == "pending":
            return step
    return None

After completing a step, update its status and attach the supporting evidence.

step = next_pending_step(plan)
if step:
    step.status = "completed"
    step.evidence.append("Inspected the relevant test file.")

This is deterministic plan management. An LLM can generate the initial plan, revise it, or recommend the next step, while the application controls status transitions.

Important: a completed plan step does not prove that the overall goal is satisfied. Completion should depend on the task's actual success criteria. "Run tests" is an action. "The required tests passed" is a verified outcome.

11. Planning strategies: reactive, plan-first, and iterative

There is no single planning strategy that works best for every task.

Strategy How it works Best for Risk
Reactive agent Chooses the next action from the latest observation. Simple tool use, short tasks, straightforward lookups. Can become shortsighted or repeat ineffective actions.
Plan-first agent Creates a sequence of steps before execution. Well-defined multi-step tasks with clear dependencies. The initial plan can become outdated when observations change.
Iterative planning agent Maintains a plan but revises it after important observations. Debugging, research, code modification, and uncertain environments. Repeated replanning can increase latency and cost without a useful progress criterion.

For our coding-agent example, iterative planning is a sensible starting point. The agent needs enough structure to make progress, but it must also respond to actual test results.

12. Error handling: what happens when a tool fails?

An agent that works only when every tool succeeds is not production-ready.

Tool failures can include:

  • Invalid arguments.
  • Missing files.
  • API rate limits.
  • Network timeouts.
  • Permission denials.
  • Malformed responses.
  • Test failures.
  • External services returning inconsistent results.

The agent runtime should classify failures rather than blindly returning the raw exception to the model.

12.1 Define structured tool results

from dataclasses import dataclass
from typing import Any
 
 
@dataclass
class ToolResult:
    ok: bool
    data: Any = None
    error_code: str | None = None
    retryable: bool = False

For example:

def safe_read_file(path: str) -> ToolResult:
    if path not in files:
        return ToolResult(ok=False, error_code="FILE_NOT_FOUND", retryable=False)
 
    return ToolResult(
        ok=True,
        data={"path": path, "content": files[path]},
    )

Now the controller can distinguish between different failure classes.

def should_retry(result: ToolResult) -> bool:
    return not result.ok and result.retryable

This is intentionally small. A production implementation should also have retry limits, exponential backoff with jitter where appropriate, per-call timeouts, and cancellation support.

12.2 Never retry every failure

Consider these examples:

Failure Typical handling
Temporary network error Retry within a bounded policy
Rate limit Respect provider guidance and back off
Invalid arguments Correct the arguments; do not repeat unchanged
Permission denied Stop or request authorization
File not found Inspect the path or report the missing file
Failed test assertion Treat it as evidence, not a transport failure
Timed-out write operation Determine whether the operation completed before retrying

Retries can be dangerous when a tool performs a side effect.

Suppose an agent submits a payment request, times out while waiting for the response, and retries the request. The original payment might already have succeeded.

For operations such as payments, deployments, and ticket creation, use idempotency keys or other mechanisms that prevent accidental duplicate execution.

A reliable agent does not merely retry. It reasons about what kind of failure occurred and what is safe to do next.

13. Termination: how does an agent know when to stop?

An agent loop needs explicit stopping conditions.

Without them, an agent can:

  • Repeat the same tool call.
  • Alternate between two unproductive actions.
  • Continue researching after the answer is sufficient.
  • Consume tokens and money without making progress.
  • Claim success without verification.

A basic controller might enforce:

MAX_STEPS = 10
MAX_TOOL_CALLS = 15

But step counts alone are insufficient. A useful agent also needs task-specific success criteria.

For a coding agent, termination might require:

  1. The intended code change exists.
  2. The relevant test command completed.
  3. The test exit code is successful.
  4. No required verification step remains outstanding.

For a research agent, success might require sufficient evidence from approved sources and an answer that clearly identifies remaining uncertainty.

13.1 Use explicit terminal states

A practical agent should distinguish between:

  • completed: the defined success criteria have been satisfied.
  • failed: the task cannot be completed under the current conditions.
  • blocked: additional permission or user input is required.
  • max_steps_reached: the execution budget was exhausted.
  • cancelled: execution was intentionally stopped.

These states are more useful than a single boolean such as done.

A final response should communicate what was actually accomplished. If the agent ran out of steps, it should not report the task as successful merely because it generated a plausible-sounding summary.

14. The ReAct pattern: reasoning and acting in a loop

One influential approach to tool-using agents is ReAct, which combines reasoning and action with observations from the environment.

A simplified conceptual trace might look like this:

Goal:         Find why the retry test fails.
 
Decision:     Inspect the retry test.
Action:       read_file("payment/tests/test_retry.py")
Observation:  The test expects validation errors not to be retried.
 
Decision:     Inspect the retry handler.
Action:       read_file("payment/retry.py")
Observation:  The handler currently retries every error.
 
Decision:     The implementation is inconsistent with the test.
Next action:  Prepare a minimal correction and run the targeted test.

The important property is that decisions are updated using observations.

The model is not merely producing a long answer about how it might solve the task. It is participating in an action–observation process.

14.1 Should you store the model's full reasoning?

Not necessarily.

An engineering system usually benefits more from storing concise, auditable information such as:

  • Action selected.
  • Arguments supplied.
  • Evidence returned.
  • Decision summary.
  • Verification result.
  • Remaining uncertainty.

These records are easier to inspect and evaluate than an unstructured transcript.

The agent's internal reasoning is not a substitute for external evidence. For example, a model stating that a test should pass does not establish that it actually passed.

14.2 ReAct is a pattern, not a complete runtime

ReAct alone does not provide:

  • Authorization.
  • Durable execution.
  • Idempotency.
  • Resource budgets.
  • Tool sandboxing.
  • Guaranteed correctness.
  • Reliable recovery after a process crash.

These are responsibilities of the surrounding agent architecture.

15. Add verification: an agent should check its own results

Consider two outcomes.

Agent A: I updated the retry logic. The issue is fixed.

Agent B: I updated the retry condition, ran the targeted test, and received exit code 0. The targeted test passed. The full integration suite has not been run.

Agent B provides a stronger engineering result because it distinguishes implementation from verification.

A verification layer can evaluate whether the claimed outcome is supported by tool evidence. For a coding workflow, the process might be:

  1. Agent proposes change
  2. Change is applied
  3. Test runner executes
  4. Exit code and output are captured
  5. Verification policy evaluates result
  6. Pass → continueFail → diagnose or revise
Verifying a code change

15.1 Example: represent test evidence

@dataclass
class TestEvidence:
    command: str
    exit_code: int | None
    output: str
    timed_out: bool = False
 
 
def tests_passed(evidence: TestEvidence) -> bool:
    return not evidence.timed_out and evidence.exit_code == 0

This function establishes a narrow condition: the command completed without a reported timeout and returned exit code zero.

It does not prove that the tests cover every relevant requirement, that the code is secure, or that the fix is correct in every environment.

A robust verification policy should define which tests must run, how their results are interpreted, and what claims can be made from those results.

16. A complete agent architecture: separating policy from runtime

At this point, we can identify a more maintainable design.

  1. User goal
  2. Task state manager
  3. Context builder
  4. LLM decision policy
  5. Action validatornot authorized → reject or request approval
  6. Tool executor
  7. Structured result
  8. State update
  9. *Verification policy*success criteria met → final result
  10. Budget available?yes → back to the context builder · no → report an incomplete task
Separating policy from runtime

This architecture separates six responsibilities.

  1. Task state manager: tracks progress, decisions, and pending work.
  2. Context builder: selects the information the model needs for its next decision.
  3. LLM decision policy: chooses an action based on the current goal and observations.
  4. Action validator and tool executor: enforce contracts and execute approved operations.
  5. Verification policy: determines whether the task's success criteria are satisfied.
  6. Execution controller: enforces budgets, cancellation, timeouts, and termination.

This separation is valuable because each component can be tested independently.

For example, you can test the action validator with invalid arguments without making an LLM API call. You can test the verification policy using recorded test results. You can evaluate the context builder using a fixed dataset.

The model then becomes one replaceable component within a controlled system.

17. What separates a prototype agent from a production agent?

The gap between a demo and a reliable agent is often less about model intelligence and more about execution engineering.

Capability Prototype Production system
Tool execution Direct function calls Validated, authorized execution
State In-memory variables Durable state where required
Error handling Basic exception handling Classified errors and bounded recovery
Execution limits Fixed loop count Step, time, token, and cost budgets
Recovery Restart from scratch Checkpoints and controlled resumption
Verification Trusts the final response Checks explicit success criteria
Security Relies heavily on instructions Enforced permissions and isolation
Observability Console logs Structured traces and operational metrics
Testing Manual trials Regression suites and failure injection
Human oversight Optional Approval gates for sensitive actions

17.1 Durable execution

Suppose an agent has completed four steps of a ten-step workflow when the application crashes.

If all state exists only in process memory, the workflow may lose its progress.

A durable runtime stores task state, action records, and important results in persistent storage. A minimal database model might include:

CREATE TABLE agent_runs (
    run_id       TEXT PRIMARY KEY,
    goal         TEXT NOT NULL,
    status       TEXT NOT NULL,
    current_step INTEGER NOT NULL DEFAULT 0,
    created_at   TIMESTAMP NOT NULL,
    updated_at   TIMESTAMP NOT NULL
);
 
CREATE TABLE agent_events (
    event_id     INTEGER PRIMARY KEY,
    run_id       TEXT NOT NULL,
    event_type   TEXT NOT NULL,
    payload_json TEXT NOT NULL,
    created_at   TIMESTAMP NOT NULL,
    FOREIGN KEY (run_id) REFERENCES agent_runs(run_id)
);

The first table stores the current execution record. The second records events such as action requests, tool results, verification outcomes, and failures.

This is only a starting schema. Production implementations also need transaction handling, concurrency control, retention policies, and protection of sensitive tool outputs.

17.2 Observability

An agent trace should help answer:

  • What was the original goal?
  • Which model and configuration were used?
  • Which tools were called?
  • How long did each call take?
  • What failed?
  • How many tokens were consumed?
  • What evidence justified the final result?
  • Why did the agent terminate?

A structured event could look like:

event = {
    "run_id": "run-1042",
    "step": 4,
    "event_type": "tool_result",
    "tool": "run_tests",
    "duration_ms": 842,
    "success": True,
    "metadata": {
        "exit_code": 0,
    },
}

Use actual measured values when generating telemetry. Avoid placing API keys, credentials, or unnecessary sensitive content into logs.

17.3 Cost and latency budgets

Every model invocation and tool call consumes resources. A useful cost model is:

C_task = Σ (i = 1 … N) [ C_model,i + C_tools,i + C_infrastructure,i ]

Here, N is the number of execution steps.

This is a conceptual accounting model; actual costs depend on provider pricing, token usage, tool infrastructure, and shared-resource allocation.

More agent steps can improve outcomes when they add useful evidence or verification. They can also waste resources when the agent repeats actions without progress.

Measure cost per successfully completed task, not just the cost of one model call.

18. When should you use an agent, and when should you avoid one?

Not every LLM application needs an agent.

A fixed pipeline is usually easier to control when the required steps are known in advance. For example:

  1. User question → Retrieve documents → Construct context → Generate answer → Validate response
A fixed pipeline

This is often preferable to an agent that repeatedly decides whether it should retrieve documents, retrieve them again, summarize them, and call additional tools.

An agent becomes more useful when the next step depends on information that has not yet been observed.

Use case Recommended starting point
Summarize a document Direct LLM call
Answer questions from a knowledge base RAG pipeline
Extract fields from invoices Structured extraction pipeline
Investigate an unknown software failure Agent with diagnostic tools
Research a topic across multiple sources Bounded research agent
Execute a predictable approval workflow Deterministic workflow
Handle a workflow with uncertain branches Agent or hybrid workflow
Perform sensitive financial operations Controlled workflow with authorization and human oversight

A hybrid approach is often the best engineering choice: use deterministic code for known processes and safety constraints, and use the LLM where interpretation or adaptive decision-making is genuinely valuable.

19. Practical exercises: build on the implementation

Use these exercises to turn the example into a more capable agent.

  1. Add a search tool (beginner). Implement search_files(query) and return matching file paths and snippets. Validate the query and cap the number of results.
  2. Add a step budget (beginner). Track model calls and tool calls separately. Stop execution when either budget is exhausted.
  3. Add persistent state (intermediate). Persist the goal, completed steps, observations, and current status in SQLite. Resume an interrupted run without repeating completed read-only actions.
  4. Add test execution (intermediate). Create a restricted test-runner tool. Capture stdout, stderr, exit code, and timeout status. Do not allow arbitrary shell commands.
  5. Add verification (advanced). Require a passing targeted test before marking a code fix verified. Report any integration tests that were not run.
  6. Add failure recovery (advanced). Simulate a temporary tool timeout, a permanent validation error, and a repeated ineffective action. Apply different recovery policies to each.

For the test-runner exercise, use a disposable project directory or isolated execution environment. A prompt telling the model not to run dangerous commands is not a substitute for enforcing what the tool can execute.

20. Final mental model

An AI agent is best understood as a controlled system that repeatedly uses a decision policy to choose actions based on a goal, current observations, and state.

The LLM may be responsible for interpreting the task and selecting the next action, but it is not the entire agent.

The surrounding runtime determines what tools exist, what actions are permitted, how state is preserved, how errors are handled, and what counts as success.

A practical architecture therefore looks like this:

  1. Goal + constraintswhat must be accomplished?
  2. Observe + build contextwhat do we know right now?
  3. Decision policywhat action should happen next?
  4. Validate + executeis the action permitted, and what actually happened?
  5. Update state + verifydid the action move us toward the goal?
  6. *Repeat*until success, failure, cancellation, or a resource limit
The agent, end to end

Conclusion

Building an agent from first principles reveals that the essential challenge is not simply getting an LLM to call a tool. It is designing a reliable loop around model-driven decisions.

The model interprets information and proposes actions. Tools interact with the environment. Observations provide feedback. State preserves progress. The controller enforces limits and permissions. Verification determines whether the intended outcome was actually achieved.

Once these responsibilities are separated, more advanced concepts become easier to understand: planning agents, coding agents, memory systems, multi-agent orchestration, durable workflows, and self-correcting systems.

The fundamental principle is simple: an LLM proposes what to do next; the agent runtime controls what happens, observes the result, and determines whether the goal has been achieved.

That is the foundation of agent engineering.

Further reading

Next in the series: Inside Tool Calling: How LLMs Use APIs, covering tool schemas, function calling, structured outputs, argument validation, tool selection, and the execution loop.