How to Review AI-Generated Code

AI coding agents fail in recognizable ways: invented packages and APIs, tests bent until they pass, changes outside the task, and summaries that claim more than the diff shows. How to review for those, in the order that catches them fastest.

7 min read

Reviewing AI-generated code means reading it in a different order from a colleague’s. Start with the task, not the code, and compare the list of changed files against what was asked. Read the test changes next, because the most common way an agent reaches “all tests pass” is by changing the tests. Then check that every new package, function and API the code calls really exists. Only then read the logic line by line, with the agent’s own summary set aside until you have formed your view. Finish by running the tests yourself and recording what you checked. The code is often fluent and well formatted, which is exactly why it needs this discipline.

The general checklist for any change, human or machine, is in the code review checklist. How to check other kinds of AI output, such as text and data, is in verifying AI-generated work. This piece is about the ways coding agents go wrong and how a reviewer catches them.

What coding agents get wrong

An agent does not get tired or skip lines, and it is usually good at syntax and at local patterns it has seen many times. Its mistakes cluster in a few places, and GitHub’s guidance on reviewing AI-generated code (opens in a new tab) names most of them: “hallucinated APIs, ignored constraints, or incorrect logic”, tests “deleted or skipped, instead of fixed”, and code that “looks right” but does not match your intent.

  • Invented dependencies and APIs. A package name that sounds right but does not exist, a method the library never had, a config option from a different version.
  • Tests bent to pass. An assertion loosened, an expected value changed to match the new output, a test marked skip, or a mock that makes the test check nothing.
  • Scope drift. Asked to fix one bug, it also renames variables, reformats a file and “improves” an unrelated function, which hides the real change in noise.
  • Ignored constraints. The rule in your instructions file, the project’s existing helper, the framework version you are on: the agent used its default instead.
  • Plausible logic that is wrong at the edges. Off-by-one, time zones, empty input, concurrent writes: the parts that need knowledge of your system rather than code in general.
  • Duplicated code. A new helper that does what an existing one already does, because the agent did not find the first one.
  • Insecure defaults. Missing authorization checks, permissive CORS, secrets in code, errors that return stack traces.
  • Summaries that overclaim. “Fixed and tested” when the tests were not run, or ran against something else.

The security side is covered in more depth in vibe coding security issues, and OWASP’s entry on improper output handling (opens in a new tab) states the principle for anything a model produces: “Treat the model as any other user, adopting a zero-trust approach.”

Review in this order

1. The task and the file list

Before any code, reread what the agent was asked to do, then look at which files changed. A bug fix in the checkout flow that touches the sign-in module, the build config and twelve test files is already telling you something. Ask for unrelated changes to be taken out; they are where regressions hide and they make the real change hard to review. This is much easier when the task itself was specific, which is the subject of giving an AI agent a task it can finish.

2. The tests

Read the test diff before the code diff. For each changed test, ask whether the expected behavior changed because the requirement changed, or because the code changed and the test was made to agree. Any deleted, skipped or loosened test needs a reason you accept. Then check that the new tests would fail without the change: the quickest way is to revert the fix locally, or break the key line on purpose, and run them. A test that stays green against broken code is decoration.

3. Dependencies and APIs

Invented package names are not a theoretical risk. A study presented at USENIX Security 2025, “We Have a Package for You!” (opens in a new tab), generated 576,000 code samples and found hallucinated packages averaging at least 5.2% for commercial models and 21.7% for open-source models, with 205,474 unique invented names. An attacker can publish a malicious package under a name models tend to invent. For every new dependency:

  • Look it up in the registry yourself. The name, author, age and download history should match what you expect.
  • Ask whether it is needed at all, or whether the standard library or an existing dependency already does it.
  • Check the lockfile changed the way you expect and nothing else was pulled in.
  • For new API calls on existing libraries, check the method exists in the version you are on, not just in the latest docs.
Quick existence checks before approving
# npm: fails with a 404 if the package does not exist
npm view <package-name> name version repository.url

# what the change added to the lockfile
git diff main -- package-lock.json | grep '"resolved"'

4. The logic

Now read the code, and read it as if nobody had summarized it. Trace the main path, then the edges: empty input, a missing record, a failed network call, two requests at once. Check every place the change touches authorization or user input. Look for a new function that duplicates an existing one. Where something looks unusual, ask why before assuming the agent knew something you did not; often it did not.

5. The claims

Last, read the agent’s summary and compare it with what you found. Anthropic’s Claude Code best practices (opens in a new tab) tell users to have Claude “show evidence rather than asserting success: the test output, the command it ran and what it returned, or a screenshot of the result.” Make that a rule in your instructions file, so every change arrives with the commit, the test command and its result. If the summary says tested and there is no output, treat it as untested.

Should another AI review it first?

A second model in a fresh session is a useful filter. The same Claude Code guide suggests a writer and reviewer pattern because a fresh context “won’t be biased toward code it just wrote”, and it also warns that a reviewer asked to find gaps “will usually report some, even when the work is sound.” Use the AI reviewer to clear the obvious problems before a person looks, tell it to report only what affects correctness or the stated requirements, and keep a person as the approval. Which tools do this is covered in AI code review tools compared.

A checklist for AI-written changes

Reviewing AI-generated code
SCOPE
[ ] I reread the task before opening the diff
[ ] Every changed file is needed for this task
[ ] No drive-by renames, reformatting or "improvements"

TESTS
[ ] Test diff read before code diff
[ ] No test deleted, skipped, loosened or re-baselined without a reason
[ ] New tests fail when the fix is reverted or broken on purpose
[ ] I ran the tests myself on the latest commit

DEPENDENCIES AND APIS
[ ] Every new package exists, is the one intended, and is needed
[ ] Lockfile changes match the new packages and nothing else
[ ] Every library method used exists in our pinned version

LOGIC AND SECURITY
[ ] Edges: empty, missing, failed call, concurrent writes
[ ] Authorization checked on new routes and queries
[ ] No secrets, no stack traces to users, no permissive CORS
[ ] No duplicate of an existing helper
[ ] Follows the rules in our instructions file

CLAIMS
[ ] Summary matches the diff
[ ] Evidence attached: commit, test command, result
[ ] What was NOT checked is written down

Recording what you checked

The review is only useful to the next person if it is written down next to the work. On a fenbs board every task has a note (the problem), a plan (how it will be done), and a test status with test notes. An AI assistant connected over MCP sets the test status with fenbs_update_item to tested, partly, failed or needs-check, and its instructions tell it never to mark something tested that it did not check. The reviewer then corrects the status to what they actually confirmed, and the change shows in History with who made it.

The reviewer’s verdict on the task
fenbs_update_item
  ref:        BUG-142
  testStatus: partly
  testNotes:  "Reviewed a91c3e0. Reverted fix: new test fails, good.
               Removed unrelated rename in utils/date.ts.
               Not checked: behavior with a real card payment."

Problems you found outside the task become their own bugs, linked to the original with relatesTo. And if the same mistake keeps coming back, such as the agent reaching for a new package when the project has a helper, record it on the Decisions and rules page as a rule. Every connected AI assistant reads the rules first, so the fix reaches every future session rather than this one.

Related

Keeping a person as the gate on merges: human in the loop for AI agents. Writing the rules an agent follows: AGENTS.md examples. Connecting an assistant to the board: Claude Code, Cursor or GitHub Copilot.

Questions people ask.

How do you review AI-generated code?

Reread the task, check which files changed, read the test changes before the code, confirm every new package and API exists, then read the logic and edge cases. Compare the agent’s summary with what you found and run the tests yourself.

What mistakes do AI coding agents make most often?

Invented packages and APIs, tests changed or skipped so they pass, edits outside the task, ignored project rules, logic that fails at the edges, duplicated helpers, insecure defaults, and summaries that claim testing that did not happen.

How do I know the tests an agent wrote are any good?

Revert the fix, or break the key line on purpose, and run them. Tests that still pass against broken code are not testing the change. Also check no existing test was deleted, skipped or loosened.

Can another AI review AI-generated code instead of a person?

It can do a useful first pass in a fresh session, but reviewers asked to find problems tend to report some even when the work is sound, and they miss what needs knowledge of your product. Keep a person as the approval.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.