An issue tracker for AI agents: filing, triage and fix
Agents now do both ends of an issue: they notice problems while working and they fix the ones they are given. That changes how an issue should be written, who decides what happens next, and why searching before filing stops being optional.
6 min read
An issue tracker for AI agents has to serve agents at both ends. Agents file issues, because an agent deep in one bug sees three others; and agents fix issues, because a well-written report is exactly the kind of instruction they follow well. That puts three demands on the tracker: every issue must be written so an agent can act on it without asking, a person must decide what gets worked on, and every filer, human or agent, must search before adding. Get those right and agents become your most diligent reporters and your fastest fixers. Get them wrong and the tracker fills with near-duplicates nobody triaged.
This is not about whether bugs and feature requests share a list; tracking bugs and feature requests in one board covers that. This is about what changes once some of the people filing and fixing are agents.
The agent as reporter, and the agent as fixer
The two jobs pull in different directions. As a reporter an agent is prolific: it notices a flaky test, a hard-coded date format, a missing null check, and if its instructions say “file anything you notice but do not fix”, it will. As a fixer it is literal: it works from what the issue says, and where the issue is vague it fills the gap with a guess.
So the rules for filing and the rules for fixing meet in the same place, which is the issue itself. An issue an agent files should be one another agent could fix. An issue a person files for an agent should be one the agent cannot misread.
What a report needs before an agent can act on it
The long-standing advice from Mozilla’s bug-writing guidelines (opens in a new tab) still holds, and it suits agents better than it ever suited people. Steps to reproduce are the most important part of a report, because a bug a developer can reproduce is very likely to be fixed. Describe what actually happened and what was expected, and keep observation separate from speculation.
For an agent, add two things: where, and how you will know it is fixed.
- Where. The file and line, the screen, the endpoint. An agent that starts at
checkout/total.ts:88spends its time fixing, not searching. - Steps to reproduce. Numbered, minimal, including setup. If the agent can run them, it can confirm the bug before touching anything.
- Expected and actual. Two short lines, facts only. “the total shows zero” is a fact; “probably a rounding bug” is a guess, and belongs lower down, labelled as one.
- Acceptance. One line that says what done means: a test that must pass, a result on screen. It is also what the reviewer checks.
Where: checkout/total.ts, applyDiscount(), around line 88 Steps: 1. Add two items to the basket on a 390px-wide screen 2. Apply code SPRING10 Expected: total drops by 10% Actual: total shows 0.00 Acceptance: checkout.spec "discount on mobile" passes Guess (unconfirmed): the discount is applied twice
Keep the problem and the plan apart
An agent that fixes an issue will learn things: the real cause, the files it had to change, the trap it nearly fell into. If it writes those into the description, the original report is lost under revisions, and the next reader cannot tell what was reported from what was discovered.
fenbs gives every task two boxes for this reason. The Problem is written once, at filing: what is wrong, why it matters, where. The Plan is how it will be done, rewritten whenever something is learnt, with History keeping the old versions. Over MCP (opens in a new tab) they are note and plan. A fixing agent writes its plan before it starts, which gives you a moment to catch a wrong approach, and comments what changed and how it was checked when it finishes.
Search before filing, every time
Mozilla’s guidelines ask reporters to check whether a bug has already been reported. With people, skipping that step produces the odd duplicate. With agents it produces a flood, because several sessions may hit the same failing test on the same morning and each will dutifully file it.
Two defences, and you want both. First, tell agents to search: on fenbs that is fenbs_search across everything the agent can see, and if the issue is there, comment on it instead of filing another. Second, have the tracker check as well, because instructions are sometimes skipped. When fenbs_create_item is called, fenbs compares the new task with the open ones and those finished in the last 14 days. If one looks like the same thing, nothing is filed and the likely matches come back with created: false.
- An open match: comment on it with what you found.
- A recently finished match: if it is the same bug back again, reopen it; if it is related but new, file it with
unlessSimilar: falseandrelatesTopointing at the old ref, so the regression is linked both ways. - Genuinely different: file again with
unlessSimilar: false. - An automated source, such as an error tracker or a CI job, can pass a
key, for example the error’s fingerprint. The same key never files twice.
fenbs_create_item {
"board": "Checkout",
"title": "Discount applied twice on mobile basket",
"kind": "bug",
"lane": "backlog",
"note": "Where: checkout/total.ts:88 ... Acceptance: checkout.spec passes"
}Triage stays with people
An agent can file, describe, link and even propose a priority in a comment. It should not decide what the team works on next. That decision weighs things an agent cannot see: a promise made to a client, a release date, the fact that the person who knows that module is on holiday.
The simplest division is by lane. Agents file into To Do. A person triages To Do: checks the kind, sets the priority, adds what is missing, and moves what should happen next into Next Up. Agents that fix take their work from Next Up and nowhere else. On fenbs, priority runs from 1 to 10 with 1 the most urgent, and a task that should not be built can be closed as a duplicate, won’t fix or cannot reproduce rather than left to rot, so the board’s numbers stay honest about what was delivered.
A small role for agents that only file
Not every agent needs to fix things. A nightly script that reads test failures, or an assistant you ask to look over a pull request, only needs to file and comment. Give it the role a human tester would get.
On fenbs that is the Reporter role: it can create tasks and comment, and edit only its own. It cannot move tasks between lanes, so it cannot promote its own reports into Next Up or close someone else’s. By default an assistant acts with the role of the person who approved it; to give it a role of its own, narrower and revocable separately, connect it from the board’s AI Assistants tab and choose the role there, as the connection guide describes. Its reports arrive with its name on them, so you can see at a glance which issues came from the nightly script and which from a customer.
Closing the loop with the reporter
When a fixing agent finishes, it comments on the issue: what changed, the commit, how it was checked. Whoever reported it, a colleague, a client, another agent, can read that on the card rather than in a chat thread. On fenbs a person can follow the bugs they care about and gets an email when one moves, which replaces most “any news on this?” messages.
Start from a template
The bug tracker template sets up a board with a Reporter role for testers and customers and sample bugs written with steps, expected and actual. For how the three kinds of task divide up, see features, enhancements and bugs. For a day-to-day rhythm with a coding agent, see the Claude Code task-tracking workflow.