How to Build Your First AI Agent, Step by Step

The smallest useful AI agent has one goal, two or three tools, a loop, a stop condition and a person who approves the risky step. Here it is in about sixty lines with the Claude Agent SDK, then how to test it and give it a task list to work from.

9 min read

To build an AI agent, give a language model one clear goal, two or three tools it can call, a loop that feeds each tool’s result back to it, a condition that makes it stop, and a point where a person approves anything that changes something real. That is the whole architecture. Everything else — memory, several agents, dashboards — is an addition you make later, once the small version works. This guide builds that small version with the Claude Agent SDK in TypeScript: an agent that reads an error log and files a bug for each new problem, asking you before it files anything.

If you want the theory of the loop first, the steps an AI agent takes to complete a task walks it step by step. If you are choosing between frameworks, agents vs tasks in multi-agent frameworks compares how they name things.

Step 1: pick one goal you can check

The goal decides everything else, so choose one where you can tell at a glance whether it worked. “Help with support” is not a goal. “For each error in today’s log seen three or more times, make sure there is a bug on the list” is: afterwards you can count the errors, count the bugs, and see which were skipped and why.

  • Small: a person could do it in under an hour.
  • Checkable: there is a right answer, or at least a list to compare against.
  • Cheap to get wrong: a wrongly filed bug costs a click to delete.
  • Repeated: you will run it again tomorrow, which is what makes building it worth it.

Step 2: give it two or three tools

A tool is a function the model can ask you to run, with a name, a description it reads to decide when to use it, and a schema for its inputs. Our agent needs three: one to read the log, one to search the existing tasks, and one to file a bug. Two of them only read. Only one changes anything, and that is the one a person will approve.

Keep tools few and their descriptions plain. Anthropic’s engineering guide to writing tools for agents (opens in a new tab) makes the point that more tools do not always lead to better outcomes, and that error messages should tell the agent what went wrong and what to try instead. A model picks tools by reading their descriptions, so a vague description gets you a confident wrong call.

Step 3: set up the project

The Claude Agent SDK (opens in a new tab) runs the same agent loop that powers Claude Code, so you do not write the loop yourself: you describe the tools and the limits, and it calls the model, runs the tools, feeds the results back and repeats. It needs Node.js 18 or later and an Anthropic API key in the ANTHROPIC_API_KEY environment variable.

Terminal
mkdir triage-agent && cd triage-agent
npm init -y
npm pkg set type=module
npm install @anthropic-ai/claude-agent-sdk zod
npm install --save-dev tsx

Step 4: write the agent

Here is the whole agent. Tools are defined with tool() and a Zod schema, as the SDK’s custom tools guide shows, and wrapped in an in-process MCP server. The task list is an array so you can run it without anything else; step 7 swaps it for a real board.

triage.ts
import { query, tool, createSdkMcpServer } from "@anthropic-ai/claude-agent-sdk";
import { z } from "zod";
import { readFile } from "node:fs/promises";
import * as readline from "node:readline/promises";

const tasks: { title: string; note: string }[] = []; // stand-in for your real task list

const readLog = tool(
  "read_error_log", "Read today's application error log.", {},
  async () => ({ content: [{ type: "text", text: await readFile("errors.log", "utf8") }] }),
  { annotations: { readOnlyHint: true } }
);

const searchTasks = tool(
  "search_tasks", "Search existing tasks for words in the title.", { words: z.string() },
  async ({ words }) => {
    const hits = tasks.filter((t) => t.title.toLowerCase().includes(words.toLowerCase()));
    return { content: [{ type: "text", text: JSON.stringify(hits) }] };
  },
  { annotations: { readOnlyHint: true } }
);

const createTask = tool(
  "create_task", "File one bug: a short title, and a note saying what, why and where.",
  { title: z.string(), note: z.string() },
  async (task) => {
    tasks.push(task);
    return { content: [{ type: "text", text: "Filed: " + task.title }] };
  }
);

const triage = createSdkMcpServer({ name: "triage", version: "1.0.0", tools: [readLog, searchTasks, createTask] });

async function ask(question: string) {
  const rl = readline.createInterface({ input: process.stdin, output: process.stdout });
  const answer = await rl.question(question);
  rl.close();
  return answer;
}

try {
  for await (const message of query({
    prompt:
      "Read today's error log. For each distinct error seen three or more times, search the tasks. " +
      "If no task matches, file one bug. Then stop and list what you filed and what you skipped.",
    options: {
      tools: [], // no built-in tools: only the three above
      mcpServers: { triage },
      allowedTools: ["mcp__triage__read_error_log", "mcp__triage__search_tasks"], // reads run freely
      canUseTool: async (name, input) => { // anything else waits for a person
        const yes = (await ask(name + " " + JSON.stringify(input) + "\nAllow? (y/n) ")) === "y";
        return yes
          ? { behavior: "allow", updatedInput: input }
          : { behavior: "deny", message: "A person declined this one. Skip it and carry on." };
      },
      maxTurns: 20, // hard stop
    },
  })) {
    if (message.type === "result") {
      console.log(message.subtype === "success" ? message.result : "Stopped: " + message.subtype);
    }
  }
} catch (error) {
  console.error("Run ended: " + error);
}

Put a few lines in errors.log and run it with npx tsx triage.ts. You will see it read the log, search for each error, and stop to ask you before every bug it wants to file.

Step 5: what each part is doing

  • The goal is the prompt. It says what done looks like — every frequent error has a task — and ends with “stop and list what you filed and what you skipped”, so the report is part of the job.
  • The tools are the three tool() calls. tools: [] removes the SDK’s built-in tools, so this agent cannot read other files or run commands; it can only do the three things you gave it.
  • The loop is query(). It calls the model, runs whatever tools it asks for, feeds the results back, and repeats until the model answers with no more tool calls.
  • The stop condition is maxTurns: 20. If the agent is still going after twenty rounds of tool calls, the run ends with a result whose subtype is error_max_turns instead of carrying on. The SDK’s agent loop guide (opens in a new tab) also documents maxBudgetUsd, a cap on spend.
  • The human check is the split between allowedTools and canUseTool. The two read-only tools are pre-approved. create_task is not, so every call to it goes to your callback, which prints what the agent wants to file and waits for y or n.

One detail in the SDK’s guide to approvals (opens in a new tab) is worth knowing before you add more tools: the callback never fires for a tool that is already approved by an allow rule or a permissive permission mode. Anything you list in allowedTools skips your check entirely. Put only tools that cannot change anything in that list.

Build it yourself or use a no-code builder?

You do not have to write code to build an agent. Workflow automation tools now offer AI steps, and model vendors offer visual builders; OpenAI, for example, built its Agent Builder as a visual canvas for building multi-step agent workflows by dragging and dropping nodes, but its deprecations page (opens in a new tab) says it was deprecated on 3 June 2026 and shuts down on 30 November 2026, pointing users to the Agents SDK or ChatGPT workspace agents instead. The choice comes down to a few questions.

  • Does the job run inside tools the builder already connects to? If so, a builder is faster to start and easier for non-developers to change.
  • Do you need a tool nobody has built, such as your own database or an internal API? Code makes that a function; a builder may make it a workaround.
  • Can you put the approval exactly where you want it, and the stop condition too? Check before you commit, not after.
  • Can you test it with a fixed set of inputs and compare the results run to run? Code makes that easy; some builders do not.
  • Where does the record of what it did live, and who can read it?

Step 6: test it before you trust it

An agent that worked once has not been tested. Keep a handful of log files where you already know the right answer — one with no frequent errors, one where every error already has a task, one with a new error, one with a malformed line — and run the agent against each after every change to the prompt or the tools. Anthropic’s engineering post on evals for AI agents (opens in a new tab) suggests starting with 20 to 50 simple tasks drawn from real failures, grading with code where you can, a model where you must, and a person for the rest, and reading the transcripts rather than only the scores. The full method is in how to evaluate AI agents.

For this agent the check is code: count the bugs filed against the errors that should have produced one. Run each case more than once, because the same input can produce a different path. Then keep a person on the output until the numbers hold up; how to review AI output without reading every line is in verifying AI-generated work, and treating a prompt change like any other change is in managing AI projects.

Step 7: give it a real task list

An array is fine for a test, but real work needs a list that outlives the run: somewhere people can see what the agent filed, change it, and give it the next job. The SDK can connect to a remote MCP server as easily as to your in-process one. A fenbs board is one: create a token for the agent by hand under Settings, with a name, the scopes it needs and an optional expiry, and pass it as a header.

Swap the array for a board
mcpServers: {
  fenbs: {
    type: "http",
    url: "https://fenbs.ai/api/mcp",
    headers: { Authorization: "Bearer " + process.env.FENBS_TOKEN },
  },
},
allowedTools: ["mcp__fenbs__fenbs_get_context", "mcp__fenbs__fenbs_search", "mcp__fenbs__fenbs_list_items"],

Now the agent reads the board’s standing notes with fenbs_get_context, searches with fenbs_search, and files with fenbs_create_item, which still goes through your approval callback because it is not in the allowed list. fenbs_create_item also refuses to file a likely duplicate of an open task and returns the match instead. Every change is recorded in the board’s History as an AI assistant acting for you, so the agent’s work can be told apart from yours. For work you want it to do rather than find, a person presses “Let AI do this” on a task with a plan, the agent takes it with fenbs_next_approved_task, and the card reads “AI done · check it” until a person confirms it. Revoking the token in Settings stops the agent at once.

Next steps

Write the tasks your agent will work from with how to write a task for an AI agent. Decide how much it may do alone with human in the loop vs fully agentic AI. Connection details and scopes are in the MCP guide and assistant tokens and scopes.

Questions people ask.

What do I need to build an AI agent?

A model you can call, one goal you can check, two or three tools written as functions, a loop that feeds tool results back to the model, a stop condition such as a maximum number of turns, and an approval step before anything that changes real data. An SDK such as the Claude Agent SDK provides the loop for you.

Can I build an AI agent without coding?

Yes. Workflow automation tools with AI steps and visual agent builders let you connect a model to apps by dragging steps together. They are quickest when the job lives inside apps the builder already supports; check that the builder itself is staying, since OpenAI’s Agent Builder shuts down on 30 November 2026. Code is better when you need your own tools, precise approval points or repeatable tests.

How do I stop an AI agent running forever?

Give it a clear finish line in the goal, and a hard limit in the code. In the Claude Agent SDK that is maxTurns, which ends the run after a set number of tool-use rounds, and maxBudgetUsd, which ends it at a spend limit.

How do I make an AI agent ask before it acts?

Pre-approve only the read-only tools and route every other tool call to an approval function. In the Claude Agent SDK that function is canUseTool, which receives the tool name and input and returns allow or deny. Tools listed in allowedTools skip it, so keep that list to tools that cannot change anything.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.