AI Agent Task Examples: 20 Tasks You Can Hand Over

Twenty real cards an AI assistant can take from a board and finish, across engineering, product, support, operations and personal work — each with what done means and whether a person should check it first.

Updated 7 min read

A good AI agent task is small enough to finish in one sitting, says where the work is, and has a line that says what done means, so both the assistant and you can tell whether it is finished. Below are twenty examples written that way, grouped by kind of work. Each has a “done means” line and a review call: yes if a person should check before it moves to Completed, no if the check is mechanical and the assistant can close it. Copy them, change the names, and you have a first week of cards.

How to read the examples

Each card is written as it would sit on a board: a title, the check that says what done means (an acceptance criterion, in Scrum terms: how that differs from a definition of done), and the review call with the reason. The reasons follow three rules. A person reviews when the result reaches someone outside the team, when it is hard to undo, or when it needs judgement the assistant does not have. Anything else — work whose check is a passing test, a count or a diff — can close without a person. Why cards should be shaped like this is covered in an AI agent task board; where people should step in is covered in human in the loop for AI agents.

Engineering

  1. Fix the failing test in checkout.spec on main. Done means: the test passes, the full suite passes, and the comment names the commit and the cause. Review: no — the suite is the check, as long as the assistant did not edit the test to make it pass.
  2. Find every place the date format is hard-coded. Done means: a comment listing each file and line, with a suggested shared helper. Review: no — it is a report; nothing changed.
  3. Upgrade the date library to the current minor version. Done means: lockfile updated, build and tests pass, changelog notes for the versions skipped are pasted in a comment. Review: yes — a person reads the changelog notes before it merges.
  4. Add a regression test for BUG-031 (double charge on retry). Done means: a test that fails on the commit before the fix and passes after, both runs shown in the comment. Review: no — the two runs are the evidence.

Product

  1. Write release notes for 1.4 from the Completed lane. Done means: a draft in the note listing every task completed since 1.3, in user language, with refs. Review: yes — customers will read it.
  2. Group this month’s feature requests by theme. Done means: a comment with each theme, the refs under it and a count. Review: no — it informs a decision a person makes afterwards; the grouping itself is easy to check at a glance.
  3. Write acceptance checks for FET-022 (export to CSV). Done means: five to ten checks in the plan, each one testable, including the empty and very large cases. Review: yes — the checks decide what counts as finished, so the owner agrees them.
  4. Find duplicates in To Do. Done means: a comment on each suspected duplicate naming the card it repeats. Nothing deleted. Review: no — it only comments; a person merges.

Support

  1. Triage the bugs filed overnight. Done means: each has a kind, a priority and a comment asking for anything missing (steps, browser, screenshot). Review: no — a person sees each one when it reaches Next Up anyway.
  2. Draft a reply to the question on BUG-044 about invoices in the wrong currency. Done means: a draft reply in a comment, with the cause if known and the ref of any fix. Review: yes — it goes to a customer.
  3. Reproduce BUG-047 (search returns deleted items). Done means: exact steps that show the fault on dev, or “cannot reproduce” with what was tried. Review: no — steps either reproduce or they do not.
  4. Turn the ten most common support questions into a help-page outline. Done means: an outline with headings and a one-line answer under each, sources linked. Review: yes — published help is read by everyone.

Operations

  1. Check every link on the marketing site. Done means: a list of broken links with the page each is on; a bug filed for each. Review: no — a link works or it does not.
  2. Summarise last week’s board activity. Done means: a comment listing what moved to Completed, what is stuck in In Progress for more than three days, and what is waiting on a person. Review: no — it reads the board and changes nothing.
  3. Draft the database migration for the new column. Done means: the script, tested on dev, with the rollback step, pasted in the plan. Run on live: never by the assistant. Review: yes — hard to undo, so a person runs it.
  4. List the connections and scheduled jobs nobody has touched in 90 days. Done means: a comment with each one, its owner and when it last ran. Review: no — it is an inventory; switching anything off is a separate card for a person.

Personal

  1. Break “plan the holiday” into tasks. Done means: six to ten tasks in To Do, one outcome each, related to the parent. Review: no — you will see each one before you do it.
  2. Compare three quotes for the boiler service from the emails I pasted. Done means: a table in a comment with price, what is included and cover period for each. Review: yes — you are about to spend money on it.
  3. Tidy my To Do lane. Done means: anything untouched for a month listed in a comment with a suggestion (delete, keep, move up). Nothing deleted. Review: no — it only suggests.
  4. Draft the thank-you email to the school fundraising team. Done means: a draft in the note, under 150 words. Review: yes — it goes out under your name.

The pattern in the review calls

Count them and twelve close without a person and eight need one. That split is typical. Most useful assistant work is reading, sorting and reporting, which is safe to close on its own evidence. The work that needs a person is a minority, and it is predictable: anything a customer, a colleague or a supplier will read, anything that spends money, anything that changes live data.

Notice also how many cards say “nothing deleted” or “a person runs it”. Splitting a job into a report the assistant can finish and an action a person takes is often the whole trick. The duplicate hunt, the stale-connection inventory and the To Do tidy all work that way. The assistant does the tedious part; the irreversible part stays with you.

Writing your own

Use the same four parts every time. It takes a minute per card, and it is the difference between an assistant that finishes and one that wanders.

A card an assistant can finish
Title:      Fix the failing test in checkout.spec on main
Kind:       bug
Note:       Fails since yesterday's merge. Card payments only.
Done means: suite passes; comment names the commit and the cause
Review:     no, unless the test itself was changed

On fenbs the note holds the problem and a separate plan holds how it will be done, which the assistant writes before it starts. When it finishes it records how the work was tested: tested, partly, failed, or “Needs owner check” for what only a person can confirm, such as a real payment or a live page. That last status is the review call made visible — the assistant says plainly that the evidence is not complete, instead of claiming a test it could not run. The four lanes stay fixed (To Do, Next Up, In Progress, Completed); review is a rule about who moves a card to Completed, not an extra lane. Kanban for AI agents shows how to set that rule.

Put these on a board

The AI assistant work log template is a board set up for exactly this kind of card. For a first pilot, how to use AI agents in project management covers which jobs to start with and how to measure them, and how to give an AI agent access to your project board covers the connection.

Questions people ask.

What is a good first task for an AI agent?

A report that changes nothing: a summary of last week’s board activity, a list of broken links, or a hunt for duplicate tasks. The output is easy to check, and a wrong answer costs nothing.

Which tasks should always be reviewed by a person?

Anything that reaches someone outside the team, anything that spends money, and anything hard to undo, such as a migration on live data or a deletion. The assistant can prepare all of these; a person should approve them.

How detailed should the “done means” check be?

One line that a stranger could check: a test that passes, a list with a count, a draft under a length. If you cannot write that line, the card is too big or too vague and should be split.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.