Types of AI Agents, With Examples
The classic answer has five types: simple reflex, model-based reflex, goal-based, utility-based and learning agents. Today’s LLM agents add a second split, between fixed workflows and agents that choose their own steps. Each type with everyday examples, and a checklist for sorting the agent in front of you.
8 min read
There are five classic types of AI agents: simple reflex agents, which react to what they see now; model-based reflex agents, which also remember what they cannot see; goal-based agents, which plan toward an end state; utility-based agents, which weigh which of several good outcomes is best; and learning agents, which get better from feedback. The list comes from Russell and Norvig’s textbook, and it still describes how any agent decides what to do. Agents built on large language models add a second, more practical split: workflows, where your code fixes the steps and the model fills them in, and agents, where the model chooses its own next step. Most real systems are a mix of both.
Where the five types come from
The taxonomy is from Stuart Russell and Peter Norvig’s Artificial Intelligence: A Modern Approach, the standard university textbook on AI. In the chapter on intelligent agents (opens in a new tab), posted on the book’s Berkeley site, an agent is “anything that can be viewed as perceiving its environment through sensors and acting upon that environment through actuators.” The authors then outline “four basic kinds of agent program,” in order of increasing generality, and explain how any of them can be turned into a learning agent. That is why some lists say four types and some say five.
For what makes software agentic in the first place, see what is agentic AI; this page sticks to the kinds.
1. Simple reflex agents
A simple reflex agent chooses its action from the current percept alone, using condition-action rules: if this, then that. It keeps no memory. The textbook’s examples are a two-square vacuum cleaner that sucks when the square is dirty and moves otherwise, and a self-driving taxi with the rule “if car-in-front-is-braking then initiate-braking.” Such agents are fast and easy to check, but they only work when the right action can be read off the current input. With partial information they can loop forever.
- A thermostat that turns on the heat below 68°F.
- An email rule that moves anything from a given sender into a folder.
- A CI check that fails the build when a lint rule is broken.
2. Model-based reflex agents
A model-based reflex agent keeps internal state: a picture of the part of the world it cannot see right now, updated with a model of how the world changes and what its own actions do. It still acts by rules, but the rules read that state, not just the latest input. In the book, the taxi keeps the previous camera frame so it can tell when brake lights come on, and keeps track of cars it cannot see when changing lanes.
- A robot vacuum that builds a map of your rooms and remembers which ones it has cleaned.
- A monitoring alert that fires only when an error rate stays high across several readings.
- A chatbot that follows a fixed script but remembers the order number you gave earlier in the conversation.
3. Goal-based agents
A goal-based agent knows what end state it wants and chooses actions by asking what will happen if it takes them. That means search and planning: looking ahead at sequences of actions rather than mapping inputs straight to outputs. The textbook calls it less efficient but more flexible. To send the taxi somewhere else you change the goal; a reflex agent would need its rules rewritten.
- A route planner finding any path from your house to the airport.
- A coding agent, such as Junie in JetBrains AI Assistant, told “make the failing test pass,” which reads code, edits and reruns the test until it does.
- A warehouse robot planning a path to a shelf.
4. Utility-based agents
Goals only say whether you got there. A utility-based agent also scores how good each outcome is, with a utility function that maps states to a number, and picks the action with the best expected utility. The book’s point is that many routes reach the destination, but some are “quicker, safer, more reliable, or cheaper than others.” Utility is what lets an agent trade off conflicting goals, such as speed against safety, and weigh likely success against importance.
- A navigation app weighing travel time against tolls and traffic.
- A thermostat schedule that balances comfort against the cost of energy at peak hours.
- An LLM system that drafts several answers and keeps the one a grader scores highest.
5. Learning agents
Any of the four can be made to learn. The textbook splits a learning agent into four parts: a performance element, which chooses actions and is what we called the whole agent before; a critic, which judges how well it is doing against a fixed performance standard; a learning element, which uses the critic’s feedback to improve the performance element; and a problem generator, which suggests new, exploratory actions so the agent finds better ones. The book’s example is the taxi learning that a quick left turn across three lanes was a bad idea from the reaction of other drivers.
- A spam filter that improves each time you mark a message as spam or not spam.
- A recommendation system that adjusts to what you actually watch.
- A game-playing program that improves by playing against itself.
Today’s LLM agents: workflows and agents
The classic types say how an agent decides. For systems built on large language models, the more useful question is who decides the next step. Anthropic’s guide to building effective agents (opens in a new tab) draws the line: workflows are systems where models and tools are “orchestrated through predefined code paths,” and agents are systems where the model dynamically directs its own process and tool use. Both start from the same building block, a model augmented with retrieval, tools and memory.
The guide names five workflow patterns. Each is explained with examples in agentic workflows, so here is one line each:
- Prompt chaining: fixed steps in sequence. Draft a product description, check it against the style guide, then shorten it.
- Routing: classify the input and send it down a matching path. Billing questions to one prompt, bug reports to another.
- Parallelization: split independent pieces and run them at once, or run the same check several times and vote.
- Orchestrator-workers: one model breaks a job it cannot predict into parts and hands them to workers, as in multi-agent workflows.
- Evaluator-optimizer: one model writes, another critiques, and the loop repeats until the critique passes.
An agent proper runs in a loop: it plans, calls tools, reads the results and decides again until the task is done or it needs a person. Anthropic’s examples are a coding agent and Claude’s computer use. Claude Code (opens in a new tab), described in its documentation as “an agentic coding tool that reads your codebase, edits files, runs commands,” is the first kind. The computer use tool (opens in a new tab) is the second: it gives Claude screenshots plus mouse and keyboard control, in a loop your application runs.
Where LLM agents fit in the classic list
- A routing workflow is close to a reflex agent: input in, category out, with the model doing the reading.
- Anything that keeps a conversation, a memory file or a plan is model-based: it carries state the current message does not show.
- A coding agent given a task is goal-based: it searches for a sequence of edits and commands that makes the tests pass.
- An evaluator-optimizer loop, or a system choosing between models by cost and quality, adds a utility function, even when the “score” is another model’s judgment.
- Learning, in the textbook sense, mostly happens outside the running agent: the model’s weights do not change during a session. Feedback that updates instructions, memory or rules between runs plays the critic’s part.
A checklist: what type is your agent?
- Does it act on the current input alone, by fixed rules? Simple reflex.
- Does it remember things the current input does not show? Model-based.
- Does it plan several steps toward an end state you gave it? Goal-based.
- Does it compare good outcomes and pick the best by a score? Utility-based.
- Does it get better from feedback over time? Learning.
- Who chooses the next step, your code or the model? Workflow or agent.
The answers tell you what to test. A reflex agent needs its rules checked; a goal-based agent needs its stopping point checked; a utility-based one needs its score checked, because it will optimize exactly what you measure. Many online lists add “hierarchical” or “multi-agent” as further types. Those describe how several agents are arranged, not how one decides, which is covered in AI agent architecture.
Giving any type of agent a goal it can read
Goal-based and utility-based agents are only as good as the goal they are given, and LLM agents forget theirs when the session ends. On fenbs, the goal lives on a task: the note says what and why, the plan says how, and the test status and test notes say how it was checked. An AI assistant connected over MCP reads the rules on the Decisions and rules page first, moves the task through To Do, Next Up, In Progress and Completed, and History records each change under its name. That is the critic’s job done by people: a record of what the agent did, to judge it by. fenbs keeps it simple, with no sprints, due dates or settable assignee.
Related
The definition: what is agentic AI and the AI agent glossary entry. Deployments companies have documented: real-world AI agent examples. Patterns in practice: agentic workflows. How an agent works through a job: steps an AI agent takes to complete a task.