Giving an AI Agent a Task It Can Finish: Scoping and Acceptance

Seven parts turn a vague request into a card an agent can finish and you can check: the outcome, where, the constraints, acceptance criteria, how it will be checked, what to do when blocked, and a size limit. With a template and three rewrites.

7 min read

A task an AI agent can finish says seven things: the outcome you want, where the work is and what the agent cannot see for itself, what it must not touch, the acceptance criteria, how the result will be checked, what to do if it gets stuck, and a size small enough for one sitting. Leave out any one and the agent fills the gap with a guess, confidently, and you find out at review. This guide takes each part in turn, then gives a template and three before-and-after rewrites.

It goes one level below the card advice in an AI agent task board and the twenty ready-made cards in AI agent task examples. Those tell you what a good card looks like. This one is about writing it.

Why one card deserves ten minutes

A colleague who gets a vague request walks over and asks. An agent rarely does. It reads the card, decides what the words most probably mean, and works until it believes it has finished. Everything you did not write becomes a decision it made for you. Ten minutes on the card is cheaper than an hour reviewing work aimed at the wrong target, and much cheaper than finding out a week later.

1. The outcome: an end state, not an activity

Start with one sentence describing the world once the work is done. “Look into the slow dashboard” is an activity; it can go on for ever and any result counts. “The dashboard loads in under two seconds on the staging data” is an outcome; it is either true or not.

  • Avoid open verbs: investigate, improve, clean up, look at, refactor. They describe effort, not results.
  • If the real job is to find something out, make the report the outcome: “A comment listing each cause of the slow load, with the query and its time.” Then a report is a finished card, not a failed fix.
  • One outcome per card. Two outcomes joined by “and” are two cards.

2. Where, and what it cannot see

Point at the place: a file and line, a screen, an endpoint, a document. Then write down the context that lives only in your head. That is the part people skip, because to them it is obvious.

  • What has already been tried, and what happened. It saves the agent repeating your afternoon.
  • Who asked for it and why. “Finance needs this for the month-end report” changes what good enough means.
  • Refs of related tasks, and any decision that shaped the area. An agent that finds odd code and “fixes” a deliberate choice has not been told it was deliberate.

Standing rules that apply to every card, such as naming, where things live or what never to deploy, do not belong on each card. Keep them in one place the agent reads first; on fenbs that is AI context, which an assistant fetches with fenbs_get_context before it starts.

3. Constraints: the edges of the work

Constraints are what the agent must not do on the way to the outcome. Without them, the shortest path wins, and the shortest path to a passing test is sometimes editing the test.

  • Scope: “Only files under src/billing. Anything else, file a new card.”
  • Interfaces: “Do not change the public API or the database schema.”
  • Dependencies: “No new packages.”
  • Evidence: “Do not modify existing tests to make them pass.”
  • Reach: “Dev only. Never run anything against live.”

4. Acceptance criteria a stranger could check

The outcome says what you want; acceptance criteria say how you will both recognise it. Write them as short statements that are each true or false, so someone who has never seen the work could tick them off.

  • The main case: “Saving a draft with no line items shows ‘Add at least one line’.”
  • Edge cases, named: empty, very large, wrong type, no permission, slow network. Agents handle the case you described and whichever others they happen to think of; list the ones you care about.
  • What must not change: “Existing invoice tests still pass. The PDF layout is untouched.” Regressions are the criteria people forget.
  • Three to seven is a good range. Fewer and something is missing; many more and the card is too big.

Ban words that cannot be checked: works well, looks good, clean, robust, fast. Replace each with a number or an observation.

5. How it will be checked, and by whom

Criteria say what; the check says how and who. Ask for evidence in a form you can read in a minute: the test names and the pass count, a before-and-after timing, a screenshot, the query it ran. And say plainly which criteria only a person can confirm, such as a real payment, an email arriving, or a page on a phone in someone’s hand.

On fenbs a task records this in its Testing field, with a status of Not tested, Tested, Partly tested, Failed or Needs owner check, and a box for what was checked, where and with what result. Over MCP these are testStatus and testNotes. Needs owner check is the honest answer when the agent did everything it could and the last check is yours; asking for it on the card beats getting a “Tested” that quietly skipped the hard part.

6. What to do when blocked

Every card should say when to stop. An agent with no stop rule will keep going past the point where a person would have asked, usually by guessing. Give it three allowed endings besides success:

  • Cannot reproduce or cannot find: stop, and comment what was tried.
  • Needs a decision: the answer depends on something only a person can choose, such as behaviour, cost or wording. Stop and ask one specific question.
  • Bigger than the card: the real problem is elsewhere or larger. Stop, file the new problem as its own card, and link the two.

On fenbs the agent comments on the card and moves it back to Next Up rather than leaving it claimed in In Progress. If the question is one the team will want on record, it can open it with fenbs_add_decision as an open decision linked to the card; a person decides it, because an assistant can write a decision down but is never the one who makes it. A new problem goes into To Do with relatesTo pointing back at the original.

7. A size limit

Aim for one sitting: one change, one area, one set of evidence. Some signs a card is too big before anyone starts it:

  • More than about seven acceptance criteria.
  • The “where” lists more than one layer: database, API and screen.
  • You can predict a decision it will need halfway through. Make that decision first, or make it a card of its own.
  • You could not review the result in fifteen minutes.

A template

On fenbs the first five parts go in the Problem box, which is written once, and the approach goes in the Plan box, which the agent fills in before it moves the card to In Progress. Over MCP they are note and plan.

Problem (note)
Outcome:     <one sentence: the end state>
Where:       <file:line, screen, endpoint>
Context:     <why, who asked, what was tried, related refs>
Constraints: <what not to touch or change>
Acceptance:
  - <main case, true or false>
  - <named edge cases>
  - <what must not change>
Check:       <evidence to paste; which items need a person>
If blocked:  <comment, move back to Next Up, ask one question>

Three rewrites

A bug

Before
Fix the export bug.
After
Outcome:     CSV export includes rows added today.
Where:       src/export/csv.ts, buildQuery(); /reports/export
Context:     Reported by support (BUG-051). Rows from today are
             missing; suspect the date filter uses UTC midnight.
Constraints: Do not change the CSV column order. No new packages.
Acceptance:
  - A row created a minute ago appears in the export
  - A row from yesterday still appears once, not twice
  - Existing export tests pass unchanged
Check:       Paste the new test name and the pass count.
If blocked:  Cannot reproduce on dev? Comment the steps tried, stop.

A feature

Before
Add dark mode.
After
Outcome:     Settings has a Light / Dark / System switch that
             changes the colour scheme on every signed-in page.
Where:       src/settings/Appearance.tsx; colours in theme.css
Context:     Several customers asked. Marketing pages are out of scope.
Constraints: Use the existing colour tokens only. No new fonts.
Acceptance:
  - The choice survives sign-out and sign-in
  - System follows the device setting
  - Text contrast meets the check we already run in CI
Check:       Screenshots of three pages in each mode.
             Needs owner check: the look on a real phone.
If blocked:  A page with hard-coded colours? List it in a comment;
             do not restyle it.

Work that is not code

Before
Sort out the supplier list.
After
Outcome:     A comment with a table of every supplier we paid this
             year: name, last invoice date, total, contract end date.
Where:       The two exports attached to this card.
Context:     For the renewal review on Friday.
Constraints: Report only. Change nothing and contact nobody.
Acceptance:
  - Every supplier in either export appears exactly once
  - Suppliers with no contract end date are flagged
Check:       State the row counts of both exports and of the table.
If blocked:  Two names that may be one supplier? List them, do not merge.

Notice that the third rewrite turns a job that sounded like an action into a report. A report the agent can finish, followed by an action a person takes, is often the safest shape for anything involving money, customers or other people’s data.

Put the card where the agent reads it

The AI assistant work log template sets up a board for cards like these. To let an assistant pull them in order, see a kanban board for Claude Code; for bugs an agent files itself, see an issue tracker for AI agents. The three kinds of card are explained in features, enhancements and bugs.

Questions people ask.

What is the most common mistake when writing a task for an AI agent?

Describing an activity instead of an outcome. “Look into the slow dashboard” has no end, so any result counts. “The dashboard loads in under two seconds on staging data” is either true or false, and both of you can tell which.

How many acceptance criteria should a card for an agent have?

Usually three to seven: the main case, the edge cases you care about by name, and what must not change. Fewer suggests something is missing; many more suggests the card should be split.

Should the agent write the plan or should I?

Either, but keep it separate from the problem. If you know the approach, write it. If not, have the agent write its plan before it starts, which gives you a moment to catch a wrong approach before any change is made.

What should an agent do when it cannot finish?

Stop, comment what it tried and what it found, and hand the card back rather than leaving it claimed. If it needs a choice only a person can make, it should ask one specific question. If it found a different or bigger problem, it should file that separately and link the two.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.