Which Tasks Can AI Agents Automate, and Which Need a Person?

Four questions sort any task on your list: can it be undone, can something other than a person check it, does it need judgement or authority, and does it touch money, people or production. The answers put it in one of three piles.

7 min read

Yes, AI agents can automate tasks, but not every task, and the dividing line is not how hard the work is. It is four things about the task itself. Can a wrong result be undone cheaply? Can the result be checked by something other than a person’s judgement, such as a test, a count or a diff? Does it need judgement or authority the agent does not hold? Does it touch money, people outside the team, or a live system? A task that is reversible and checkable, and needs neither judgement nor money, can be automated. A task that fails one of those can still be prepared by an agent, with a person making the last move. A few stay with a person altogether.

This post is the sorting method, for your own list. If you want ready-made cards, AI agent task examples has twenty. Where exactly a person steps in once a task is in the middle pile is in human in the loop for AI agents, and how to check what comes back is in verifying AI-generated work.

The four questions

  1. Reversible: if the agent gets it wrong, can we undo it ourselves, quickly, and completely? A reverted commit, a card moved back and a soft-deleted record are reversible. A sent email, a payment and a dropped table are not.
  2. Checkable: can something other than a person’s opinion confirm it is right? A passing test suite, a row count that reconciles, a link that loads and a schema that validates are checks. “Does this read well to a client?” is not.
  3. Judgement or authority: does the task need a decision the agent is not entitled to make? Choosing what to cut, setting a price, agreeing to a contract term or deciding about a colleague are decisions somebody answers for.
  4. Money, people, production: does the result spend money, reach someone outside the team, or change a live system that others depend on?

The second question matters more than it looks. Anthropic’s guidance on building effective agents (opens in a new tab) says agents need “ground truth” from the environment at each step, such as tool results or code execution, and gives coding as a good fit because solutions are verifiable through automated tests. An agent working without a check is guessing and telling you it is sure.

Three piles

  • Automate: reversible, checkable, no judgement or authority needed, no money, outsiders or live systems. The agent does it and records what it did.
  • Agent prepares, person decides: fails exactly one or two of the questions, usually reversibility or the money-and-people test. The agent does everything up to the last step and stops.
  • Keep with a person: needs authority the agent cannot hold, or judgement about a person. The agent may gather facts; the decision and the act stay human.

Write the four answers next to each task, not a gut feeling about the pile. Two people sorting the same list should land in the same place, and when they do not, the disagreement is usually about one question, which is quick to settle.

A sorting sheet
Task                               Rev  Check  Judg  $/ppl/prod  Pile
Rename a function across the repo  yes  tests  no    no          automate
Import 300 rows of tasks from CSV  yes  count  no    no          automate
Reply to a complaint               no   no     yes   people      prepare
Run the migration on live          no   dev    no    prod        prepare
Decide who leads the next project  -    no     yes   people      person

Worked examples: automate

  • Renaming a function everywhere it is used. Reversible with a revert, checked by the build and the tests, no decision in it. Automate, on a branch.
  • Importing a spreadsheet of tasks into a board. Reversible by deleting what was added, checked by counting rows in and items out. Automate, and have the agent report both numbers.
  • Tagging last month’s support tickets by product area. Reversible, and a sample of twenty is easy to check by eye. Automate; spot-check a sample once.
  • Finding every page on the site that still mentions an old plan name. It changes nothing, so it cannot go wrong in a way that matters. Automate; the fixes are separate tasks.

Worked examples: agent prepares, person decides

  • Replying to a customer complaint. Cannot be unsent and reaches someone outside. The agent reads the history and drafts; a person edits and sends.
  • Running a database migration on live data. Checkable on a copy, but not reversible once it has run on live. The agent writes the script and the rollback and tests both on a development copy; a person runs it.
  • Issuing a refund above a set amount. Money. The agent checks the order against the policy and proposes the figure with its reasons; a person approves.
  • Publishing a help article. Reaches every reader, and accuracy is partly judgement. The agent drafts and links its sources; a person reads and publishes. Human in the loop for generative AI content covers that review.
  • Changing who has access to a system. Hard to notice when wrong, and it moves authority. The agent prepares the change; a person makes it.

OWASP puts the same rule in security terms. Its entry on excessive agency (opens in a new tab) traces damaging agent actions to too much functionality, too many permissions or too much autonomy, and recommends human-in-the-loop control so that a person approves high-impact actions before they are taken. Its example is an assistant that can send email when it only needed to read it.

Worked examples: keep with a person

  • Deciding who leads the next project, or how someone’s work is rated. Judgement about a person, and one they will hold someone to account for. An agent may summarise the facts; it should not rank people. In the EU, AI used to allocate tasks on personal traits or to evaluate performance is listed as high-risk, as human oversight under the EU AI Act explains.
  • Choosing what to cut from a release. The trade-off is the job, and the person who makes it answers for it.
  • Agreeing a price, a contract term or a settlement. Authority that belongs to a named role, not to whoever is holding the keyboard.
  • Answering a regulator, a lawyer or the press. The agent can collect the record; the words carry the organisation’s name.

Tasks that change pile

The same verb can sit in different piles depending on where it happens. Sending an email to your own team is reversible enough; sending one to every customer is not. Deploying a preview is automatable; deploying to production is not. Deleting a record that can be restored is fine; a hard delete is not. So sort the task as it will actually run, in the environment it will run in, not the kind of task in general.

When a task lands in the middle pile, try splitting it. “Clean up duplicate customers” becomes “list suspected duplicates with the reason for each”, which is automatable, and “merge the confirmed ones”, which a person does. The tedious part moves to the agent; the irreversible part stays where it was. Most middle-pile work splits this way, and the split is where most of the time saving comes from.

Enforce the piles in the tools, not the prompt

An instruction in a prompt is a request the agent may forget. A permission is a rule. Claude Code, for example, has permission rules (opens in a new tab) in three lists, allow, ask and deny, evaluated in the order deny, then ask, then allow, so a command you put on the ask list prompts you even when a broader allow rule also matches. The automate pile goes on allow, the middle pile on ask, and anything that should never happen from that session on deny.

On a fenbs board the same sort is made from what the board already has. An assistant connects with the scopes you tick, read, write and comment, so one that should only report can be given read and comment and nothing else. Roles are defined per company, and moving a card between lanes is its own permission, so an assistant can add and edit tasks while a person decides what reaches Completed. When the assistant finishes something only a person can confirm, such as a real payment, it sets the task’s testing status to Needs owner check, and those tasks stand out until somebody acts. Every change is recorded in History with who made it, an assistant’s as “Claude via” the person it acts for.

Re-sort on a schedule

The piles are not permanent. A task moves towards automate when a check appears for it: once the migration has a dry run that proves it on a copy of live data, or the reply has a template approved for routine cases. It moves the other way when something changes the stakes, such as a new customer contract or a system others now depend on. Look at the sheet once a quarter, and after anything went wrong, and move tasks one pile at a time, in writing.

Related

Put the sorted tasks on a board with the AI assistant work log template, connect an assistant with only the scopes it needs using the connection guide, and see assistant tokens and scopes for how scopes and roles combine. For a first pilot, read how to use AI agents in project management.

Questions people ask.

Can AI agents automate tasks completely?

Some, yes: tasks that can be undone cheaply, that something mechanical such as a test or a count can check, and that need no judgement, money, outside audience or live system. Many more can be prepared by an agent up to the last step, with a person making the final move.

Which tasks should an AI agent never do on its own?

Decisions that need authority or judgement about people, such as rating someone’s work, agreeing a price or answering a regulator. Also any irreversible act, such as a payment, a customer email or a change to live data, unless a person has approved that specific action.

What if a task is half automatable?

Split it. Turn the task into a report the agent can finish, such as a list of suspected duplicates with reasons, and an action a person takes, such as merging the confirmed ones. Most tasks in the middle split this way.

How often should I re-sort tasks?

Once a quarter and after any mistake. Move a task towards automation when a reliable check appears for it, and back towards a person when the stakes change, one pile at a time and in writing.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.