The Steps an AI Agent Takes to Complete a Task
Read the task, plan, choose a tool, act, look at what happened, check it against the goal, go round again, and report. What happens at each step, how each one fails, and where a person should step in.
7 min read
An AI agent completes a task by going round a loop. It reads the task and whatever context it has been given, makes a plan, picks a tool, uses it, looks at what came back, checks that result against the goal, and either goes round again or stops and reports. Anthropic’s engineering guide to building effective agents (opens in a new tab) puts it in one line: agents are typically just language models using tools based on environmental feedback in a loop. Everything interesting about an agent, good and bad, happens at one of those steps, so it pays to know them by name.
This post walks the loop in order. For each step: what the agent is doing, the way it usually goes wrong, and what a person can do about it. If you want the definition of an agent first, the glossary has what an AI agent is.
The loop at a glance
- Read the task and the context around it.
- Plan: decide what done looks like and the first few moves.
- Choose a tool: search, read a file, call an API, run a command.
- Act: make the call.
- Observe: read what came back, including errors.
- Check: compare where things stand with what done looks like.
- Repeat or stop: go round again, ask for help, or finish.
- Report: say what was done, what was checked and what was not.
The idea of interleaving thinking and doing is older than most agent products. The 2022 ReAct paper (opens in a new tab) by Yao and colleagues showed a model producing reasoning traces and actions in turn, with the reasoning used to form, track and update a plan and to handle exceptions, and the actions used to fetch information from outside. The steps above are that pattern, with a start and an end added.
1. Read the task and the context
The agent knows only what it is given and what it can look up. A task that says “fix checkout” gives it a direction; a task that says what is broken, where, and how to tell it is fixed gives it a destination. The context around the task matters as much: the conventions of the codebase, what must never be touched, what was decided last month.
Where it fails: a vague task, or context that is out of date. The agent does not know what it was not told, and it will fill the gap with something plausible. Where a person comes in: before the loop starts, by writing the task well. Giving an AI agent a task it can finish covers the outcome, the constraints and the acceptance check. On a fenbs board an assistant reads the task with fenbs_get_item and the board’s standing notes with fenbs_get_context, so the context is written once and read every time.
2. Plan
Before acting, a capable agent sketches the route: the goal restated, the first moves, how it will know it has arrived. Some tools make this a separate stage you can see and approve; in Claude Code that is plan mode, where it explores and writes a plan without editing anything until you accept it.
Where it fails: the plan is too big for one sitting, or it solves a slightly different problem from the one you had. Both are cheaper to catch here than at the end. Where a person comes in: approve the plan for anything hard to undo, and split work that will not fit. AI agent task decomposition is about making the pieces small enough.
3. Choose a tool
The agent has a list of tools, each with a name, a description and the inputs it takes. It picks one by matching what it needs next against those descriptions. This is the step people see least and blame most. A tool with a vague description, or two tools with near-identical names, leads to the wrong call made with total confidence.
Where a person comes in: mostly as whoever builds or connects the tools. Fewer, clearer tools beat many overlapping ones. When an assistant keeps choosing the wrong one, how to debug MCP tools shows how to see the tool the way the model sees it.
4. Act
The agent makes the call: edits the file, runs the test, creates the task, sends the request. This is the only step with side effects, which is why every safeguard sits here. There are two gates, and they do different jobs. The client can ask you before it runs a tool. The system on the other end decides whether the call is allowed at all, based on who the agent is acting as.
Where it fails: an action that should have been a question, or a permission wider than the job. Where a person comes in: by setting both gates. Keep confirmations on for writes, and connect the agent with the least access that does the work. On fenbs an assistant acts as the person who connected it, narrowed by the scopes ticked at sign-in, so it can never do more than that person could.
5. Observe
After each action the agent reads the result: the test output, the file it just wrote, the error. Anthropic’s guide calls this getting ground truth from the environment at each step, and it is what separates an agent from a model guessing in the dark. A good result lets it move on. A clear error lets it correct itself.
Where it fails: an error the agent cannot read, or one it reads and ignores. A refusal that says only “forbidden” invites retries; one that says which permission is missing and which role the agent holds gives it something to report. fenbs returns refusals as that kind of sentence, as a tool result the assistant can relay, not a broken call.
6. Check
Checking is observing with the goal in mind: not “did the command run?” but “is the thing now true that the task said should be true?”. Anthropic’s documentation for Claude Code describes its agentic loop (opens in a new tab) as three phases that blend together — gather context, take action, verify results — with a bug fix cycling through all three repeatedly.
Where it fails: the agent checks something easier than the goal, or declares success without checking at all. Where a person comes in: by putting the check in the task (“the test in checkout.spec.ts passes”) so there is one right answer, and by reviewing the evidence rather than the claim. Verifying AI-generated work sets out a review routine that scales.
7. Repeat or stop
If the check fails, the agent goes round again with what it learnt. If it passes, it stops. Frameworks give the loop a hard edge too: OpenAI’s guide to running agents (opens in a new tab) describes a runner that calls the model, runs any tool calls it produced and continues, and returns once the model gives a final answer with no more tool work, with limits on how many turns a run may take. Anthropic’s guide likewise recommends stopping conditions such as a maximum number of iterations, and pausing for human feedback at checkpoints or blockers.
Where it fails: loops that keep trying the same fix, and agents that give up quietly and describe the attempt as a result. Where a person comes in: interrupt a loop that is going nowhere, and tell the agent in advance what to do when blocked, which is usually “stop and say why”. When a Claude Code task is stuck lists the common causes.
8. Report
The last step is the one most setups skip. The agent should say what it changed, how it checked it, and what it did not check, somewhere that outlives the chat window. A transcript is a poor report: long, mostly tool output, and gone when the session closes.
On a board, the report is the task itself. A fenbs task has a Testing part with a status (Not tested, Tested, Partly tested, Failed, Needs owner check) and a box saying what was checked; the assistant comments what changed and moves the task. A task a person pre-approved for AI shows “AI done · check it” until a person confirms it, which puts the final check where it belongs.
Where a person steps in, in one list
- Before step 1: write the task with an outcome and a check.
- At step 2: approve the plan for anything hard to undo.
- At step 4: keep confirmations on for writes, and keep access narrow.
- At step 7: interrupt loops, and set what happens when the agent is blocked.
- After step 8: check the evidence before the work counts as done.
None of these slows a good agent down much, and each one catches a failure at the step where it is cheapest to fix. Human in the loop for AI agents goes further into how many of these checkpoints a team needs.
Give the loop somewhere to start and finish
A board gives an agent its task at step 1 and a place to report at step 8. See how it works, connect an assistant with the MCP guide, or read agentic workflows explained for how several loops combine into a larger job.