Verifying AI-Generated Work: A Review Process That Scales
Reading everything an agent produces does not scale, and reading nothing is not review. A process that holds up has three parts: evidence the agent must attach, a check that fits the kind of output, and a verdict written down where the next person will see it.
7 min read
A human in the loop verification process for AI-generated output has five parts. The agent attaches evidence with every piece of finished work: the commit, the command it ran and what came back, the sources it used. The reviewer checks against a short list written for the kind of output, because code, text, data and decisions go wrong in different ways. Irreversible or external work gets a full review; frequent, reversible work gets sampled. The verdict is recorded on the work item, with what was checked and what was not. And anything the reviewer cannot confirm is escalated to a named person rather than passed. Done that way, review time grows with risk, not with volume.
Why review is the bottleneck
Agents shorten the doing. They do not shorten the checking, and they make it easier to skip. OWASP’s Top 10 for LLM applications (opens in a new tab) lists misinformation as LLM09:2025 and describes the human side of it plainly: “Overreliance occurs when users place excessive trust in LLM-generated content, failing to verify its accuracy.” Its mitigations include human oversight and fact-checking “especially for critical or sensitive information”, and it notes that models can suggest “insecure or non-existent code libraries”.
The answer is not to read harder. It is to make each review cheaper, and to spend full reviews only where they pay. The rest of the process follows from that.
Step 1: the agent attaches evidence
An agent’s “done” is a claim. Evidence turns it into something a reviewer can check in a minute rather than reproduce in an hour. Claude Code’s own best-practice guide (opens in a new tab) says the same: “Have Claude show evidence rather than asserting success: the test output, the command it ran and what it returned, or a screenshot of the result.” Write the requirement into the agent’s rules file so it happens every time.
- Code: the commit hash, the test command, and the summary line of its output. Name any test that was added, changed or skipped.
- Text: a source link for every factual claim, and a note of anything it could not source and left out.
- Data: row counts before and after, the query or script used, and a handful of changed records listed by id.
- Decisions and recommendations: the options considered, the reason for the choice, and what would change the answer.
- Anything visual: a screenshot of the result, taken after the change.
Ask for what was not checked as well. “Tests pass; not tried on mobile” is far more useful than “tests pass”, because it tells the reviewer where to look.
Step 2: check by output type
One generic checklist ends up checking nothing. Keep a short list per type of output, and keep each list short enough to be used every time.
Code
- Does the diff stay inside the area the work item named?
- Were tests edited, weakened or skipped to make them pass?
- Are new dependencies real, maintained and needed? Hallucinated package names are a known attack route.
- Does the commit exist, and does it match the one quoted?
Text
- Follow two or three source links and confirm they say what the text claims.
- Check every number, name, date and quotation against its source.
- Read it once for tone and audience: fluent and wrong reads exactly like fluent and right.
Data
- Do the before and after counts reconcile with what was meant to change?
- Open a few of the listed records and check them by hand.
- Look for the edge cases: nulls, duplicates, records that matched a pattern by accident.
Decisions
- Is the reasoning based on facts you can check, or on assumptions stated as facts?
- Were the obvious alternatives considered and fairly described?
- Is it a decision a person should own? If so, it is a recommendation until one does.
Step 3: full review or sample
Not every output earns a full review. Decide per kind of work, not per item, so nobody has to judge afresh each time.
- Full review, every time: anything hard to undo, anything that reaches customers or the public, anything touching money, access or live data, and anything a mechanical check cannot cover.
- Sample: frequent, reversible work with a mechanical check already in place, such as lint fixes, label changes, filed bugs and routine refactors behind passing tests.
- No review beyond the record: reading, searching, drafting that nobody will act on unseen.
For sampled work, pick a fixed number from each agent each week, say five, chosen at random rather than the ones that look interesting. If one fails, widen the sample for that kind of work until the pattern is understood; if several weeks pass cleanly, the number can stay small. Sampling tells you whether a kind of work can be trusted. It does not catch every error, so reserve it for work where an occasional error is cheap and reversible. Which jobs belong in which group is the subject of human in the loop vs fully agentic AI.
A second agent can take some of the load. Claude Code’s guide describes a writer and reviewer pattern with a fresh session, since a fresh context “won’t be biased toward code it just wrote”, and a bundled /code-review skill that reviews the current diff in a separate subagent. Treat a second model as a filter that makes human review quicker, not as the review itself.
Step 4: record the verdict
A review nobody can see did not happen, as far as the next person is concerned. Record the verdict on the work item itself, in words, with what was checked and what was not.
On fenbs every task has a Testing section: a status and a notes box. The statuses are Not tested, Tested, Partly tested, Failed, and Needs owner check for what only a person can confirm, such as a store build or a real payment. Over MCP they are testStatus and testNotes, and the tool description tells assistants never to claim tested for a check they did not run. A change of status is a line in History, recorded with who made it.
fenbs_update_item
ref: BUG-031
testStatus: partly
testNotes: "Unit tests 41/41 on 7f3e21a. Retry path checked
on a dev copy. Not checked: a real card payment."The reviewer then does the part the assistant could not, updates the status, and moves the card to Completed. Moving between lanes is its own permission on fenbs, so a team can keep every move to Completed with a person. Before a deploy, the board’s filter for “Completed, not tested” answers the one question that matters: what is about to ship that nobody checked?
Step 5: escalate what you cannot confirm
Reviewers get stuck, and a stuck reviewer tends to pass work rather than hold it. Give them three explicit ways out instead.
- Cannot confirm it yourself: mark it Needs owner check and name who can, in a comment. Those tasks stand out on the board until somebody acts.
- Checked and wrong: mark it Failed, say what failed, and move it back to In Progress. If the fault is a separate problem, file it as a bug linked with
relatesTo. - Not sure it should exist at all: comment the doubt and stop. Scope and priorities are a person’s call, as human in the loop for AI agents explains.
Keeping the process honest
Once a month, read a handful of tasks marked Tested and redo one check yourself. If the notes said more than was done, tighten the evidence rule. The wider monthly review of who can do what is in how to audit AI agents; the per-item record it relies on is the one above.
Related
Write work items an agent can finish and prove with giving an AI agent a task it can finish. See how the Completed lane stays meaningful in kanban for AI agents, and read what History records in an audit trail for AI agents.