An issue tracker for AI agents: filing, triage and fix

Agents now do both ends of an issue: they notice problems while working and they fix the ones they are given. That changes how an issue should be written, who decides what happens next, and why searching before filing stops being optional.

6 min read

An issue tracker for AI agents has to serve agents at both ends. Agents file issues, because an agent deep in one bug sees three others; and agents fix issues, because a well-written report is exactly the kind of instruction they follow well. That puts three demands on the tracker: every issue must be written so an agent can act on it without asking, a person must decide what gets worked on, and every filer, human or agent, must search before adding. Get those right and agents become your most diligent reporters and your fastest fixers. Get them wrong and the tracker fills with near-duplicates nobody triaged.

This is not about whether bugs and feature requests share a list; tracking bugs and feature requests in one board covers that. This is about what changes once some of the people filing and fixing are agents.

The agent as reporter, and the agent as fixer

The two jobs pull in different directions. As a reporter an agent is prolific: it notices a flaky test, a hard-coded date format, a missing null check, and if its instructions say “file anything you notice but do not fix”, it will. As a fixer it is literal: it works from what the issue says, and where the issue is vague it fills the gap with a guess.

So the rules for filing and the rules for fixing meet in the same place, which is the issue itself. An issue an agent files should be one another agent could fix. An issue a person files for an agent should be one the agent cannot misread.

What a report needs before an agent can act on it

The long-standing advice from Mozilla’s bug-writing guidelines (opens in a new tab) still holds, and it suits agents better than it ever suited people. Steps to reproduce are the most important part of a report, because a bug a developer can reproduce is very likely to be fixed. Describe what actually happened and what was expected, and keep observation separate from speculation.

For an agent, add two things: where, and how you will know it is fixed.

  • Where. The file and line, the screen, the endpoint. An agent that starts at checkout/total.ts:88 spends its time fixing, not searching.
  • Steps to reproduce. Numbered, minimal, including setup. If the agent can run them, it can confirm the bug before touching anything.
  • Expected and actual. Two short lines, facts only. “the total shows zero” is a fact; “probably a rounding bug” is a guess, and belongs lower down, labelled as one.
  • Acceptance. One line that says what done means: a test that must pass, a result on screen. It is also what the reviewer checks.
A bug an agent can pick up
Where: checkout/total.ts, applyDiscount(), around line 88
Steps:
  1. Add two items to the basket on a 390px-wide screen
  2. Apply code SPRING10
Expected: total drops by 10%
Actual: total shows 0.00
Acceptance: checkout.spec "discount on mobile" passes
Guess (unconfirmed): the discount is applied twice

Keep the problem and the plan apart

An agent that fixes an issue will learn things: the real cause, the files it had to change, the trap it nearly fell into. If it writes those into the description, the original report is lost under revisions, and the next reader cannot tell what was reported from what was discovered.

fenbs gives every task two boxes for this reason. The Problem is written once, at filing: what is wrong, why it matters, where. The Plan is how it will be done, rewritten whenever something is learnt, with History keeping the old versions. Over MCP (opens in a new tab) they are note and plan. A fixing agent writes its plan before it starts, which gives you a moment to catch a wrong approach, and comments what changed and how it was checked when it finishes.

Search before filing, every time

Mozilla’s guidelines ask reporters to check whether a bug has already been reported. With people, skipping that step produces the odd duplicate. With agents it produces a flood, because several sessions may hit the same failing test on the same morning and each will dutifully file it.

Two defences, and you want both. First, tell agents to search: on fenbs that is fenbs_search across everything the agent can see, and if the issue is there, comment on it instead of filing another. Second, have the tracker check as well, because instructions are sometimes skipped. When fenbs_create_item is called, fenbs compares the new task with the open ones and those finished in the last 14 days. If one looks like the same thing, nothing is filed and the likely matches come back with created: false.

  • An open match: comment on it with what you found.
  • A recently finished match: if it is the same bug back again, reopen it; if it is related but new, file it with unlessSimilar: false and relatesTo pointing at the old ref, so the regression is linked both ways.
  • Genuinely different: file again with unlessSimilar: false.
  • An automated source, such as an error tracker or a CI job, can pass a key, for example the error’s fingerprint. The same key never files twice.
An agent filing a bug it noticed
fenbs_create_item {
  "board": "Checkout",
  "title": "Discount applied twice on mobile basket",
  "kind": "bug",
  "lane": "backlog",
  "note": "Where: checkout/total.ts:88 ... Acceptance: checkout.spec passes"
}

Triage stays with people

An agent can file, describe, link and even propose a priority in a comment. It should not decide what the team works on next. That decision weighs things an agent cannot see: a promise made to a client, a release date, the fact that the person who knows that module is on holiday.

The simplest division is by lane. Agents file into To Do. A person triages To Do: checks the kind, sets the priority, adds what is missing, and moves what should happen next into Next Up. Agents that fix take their work from Next Up and nowhere else. On fenbs, priority runs from 1 to 10 with 1 the most urgent, and a task that should not be built can be closed as a duplicate, won’t fix or cannot reproduce rather than left to rot, so the board’s numbers stay honest about what was delivered.

A small role for agents that only file

Not every agent needs to fix things. A nightly script that reads test failures, or an assistant you ask to look over a pull request, only needs to file and comment. Give it the role a human tester would get.

On fenbs that is the Reporter role: it can create tasks and comment, and edit only its own. It cannot move tasks between lanes, so it cannot promote its own reports into Next Up or close someone else’s. By default an assistant acts with the role of the person who approved it; to give it a role of its own, narrower and revocable separately, connect it from the board’s AI Assistants tab and choose the role there, as the connection guide describes. Its reports arrive with its name on them, so you can see at a glance which issues came from the nightly script and which from a customer.

Closing the loop with the reporter

When a fixing agent finishes, it comments on the issue: what changed, the commit, how it was checked. Whoever reported it, a colleague, a client, another agent, can read that on the card rather than in a chat thread. On fenbs a person can follow the bugs they care about and gets an email when one moves, which replaces most “any news on this?” messages.

Start from a template

The bug tracker template sets up a board with a Reporter role for testers and customers and sample bugs written with steps, expected and actual. For how the three kinds of task divide up, see features, enhancements and bugs. For a day-to-day rhythm with a coding agent, see the Claude Code task-tracking workflow.

Questions people ask.

Should AI agents file their own bug reports?

Yes, if they search first and file into a lane a person triages. An agent working on one problem often notices others, and a filed report with steps and a location is worth far more than a remark lost in a transcript.

How do I stop agents filing duplicate issues?

Tell them to search before filing and to comment on an existing issue instead. On fenbs the tracker also checks: fenbs_create_item holds back a task that looks like an open one or one finished in the last 14 days and returns the likely matches, and an automated source can pass a key so the same event never files twice.

What should a bug report for an AI agent contain?

Where the problem is, numbered steps to reproduce, what was expected and what actually happened, and one line saying how you will know it is fixed. Keep guesses about the cause separate and labelled as guesses.

Should an agent decide which issues get fixed next?

No. It can propose a priority in a comment, but a person should triage: set the priority and move what is next into Next Up. Fixing agents then take their work from Next Up only.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.