When an AI Agent Makes a Mess: An Incident Checklist

An agent deleted the wrong thing, pushed what it should not have, or read something it was never meant to see. Contain, assess, restore, communicate and learn, in that order, mapped to NIST’s current incident response guidance.

8 min read

AI agent incident response has five steps, in this order. Contain: stop the session and revoke the agent’s own credentials, not the person’s, then rotate any secret it could have read. Assess: collect the records from every system it touched, before they roll over, and work out what it read, what it changed and where it could have sent data. Restore: undo what can be undone from those records, checking each backup before you use it. Communicate: tell the owner of each affected system, and anyone outside if data left. Learn: narrow the access that made it possible and write the change down. The first two steps usually take under an hour and decide how bad the rest is.

The short version of this plan is one of the practices in AI agent security best practices. This is the long version, for the day you need it.

Three kinds of agent incident

  • A mistake. The agent did what it understood you to ask: closed forty tasks, deleted the “duplicate” branch, rewrote a migration. Nobody attacked anything, but the damage is real.
  • A manipulation. The agent read something written by someone else, such as an issue, an email or a web page, and followed instructions in it. The actions may look like a mistake; the cause is outside.
  • An exposure. A secret, customer record or private file was read, printed, committed or sent somewhere it should not be. Nothing may look broken at all.

You rarely know which one you have at the start, so treat every incident as if it might be the second or third until the records say otherwise. Containing a mistake as if it were an attack costs a few minutes. The reverse can cost much more.

How NIST frames it now

NIST’s incident response guidance was rewritten in April 2025 as SP 800-61 Revision 3 (opens in a new tab), which replaces the 2012 edition. It drops the old four-phase cycle and maps incident response onto the six functions of the Cybersecurity Framework 2.0. Govern, Identify and Protect are preparation: broader risk management that supports response. Detect, Respond and Recover are the response itself. Lessons learned sit in the Improvement category of Identify, and NIST’s point is that they should be shared as soon as they are found, not saved for a meeting after recovery. The steps below name the categories they fall under, which helps if a customer or insurer asks how you handle incidents.

1. Contain: stop it and take its keys (RS.MI)

NIST’s Incident Mitigation category covers containment and eradication: stopping the incident from spreading. For an agent that means ending what is running and removing what lets it act again.

  1. Stop the session. Interrupt the agent, cancel a cloud agent’s run, disable a scheduled job. How to do that cleanly in Claude Code is in how to stop a Claude Code task midway.
  2. Revoke the agent’s credentials in every system it held one, not the person’s login. If the agent had its own token, the person can keep working while you investigate. If it shared the person’s login, you now have to sign them out as well, which is the best argument for giving agents their own tokens.
  3. Look for credentials the agent created. GitHub notes that when an organisation owner revokes a fine-grained personal access token (opens in a new tab), SSH keys created by that token keep working. Keys, webhooks, deploy keys and new service accounts made during the session need removing separately.
  4. Rotate every secret the agent could have read, whether or not you think it did. That includes anything in an environment file, a shell variable or a config file in its reach.

If a secret was committed, rotate it before you clean it out of the repository. GitHub’s guidance on remediating a leaked secret (opens in a new tab) is explicit that removing it from the code is not enough, that revoking it with its provider is the most important step, and that rewriting history is destructive and should be weighed carefully.

2. Assess: read the records before they go (RS.AN)

Incident Analysis is about investigating enough to respond well. Collect first, interpret second, because some records roll over.

  • Version control: the commits, branches and pushes from the session, and any force pushes. On a local machine, git reflog shows where branches pointed before they moved.
  • The agent’s own session: the transcript, and any telemetry you export from it.
  • Each tool’s history: the board, the help desk, the document store, filtered to the agent’s name.
  • Provider and platform logs: your code host’s security and audit logs, CI runs, cloud access logs, and the provider logs for any secret that was exposed.

Then build a short timeline and answer three questions. What did it read, especially from outside the team, in the minutes before things went wrong? What did it change, and where? Where could it have sent data, given its network access and its tools? The answer to the first tells you whether this was a manipulation. The answer to the third tells you whether you also have an exposure.

3. Restore: undo from the records (RC.RP)

NIST’s Incident Recovery Plan Execution category includes checking the integrity of backups before restoring from them. For agent incidents, the same idea applies to every undo button: know what it covers before you press it.

  • Code: revert the agent’s commits on a branch and merge the revert through your usual review, rather than resetting shared history.
  • Session edits: Claude Code’s checkpoints (opens in a new tab) can rewind file edits made by its own editing tools with /rewind. The documentation is clear about the limits: changes made by shell commands are not tracked, most subagent edits are not restored, and checkpoints are not a replacement for version control.
  • Data: restore from a backup taken before the session started, after checking it is complete.
  • Tools: use each tool’s own undo, and check what it keeps and what it does not.
Terminal: revert an agent’s commits without rewriting history
git log --oneline --since="2 hours ago"   # find the agent's commits
git revert --no-edit <first>^..<last>     # one revert per commit, newest first
git push origin HEAD                      # then review and merge as usual

4. Communicate: who needs to know (RS.CO, RC.CO)

NIST separates communication during the response from communication during recovery, and both are about coordinating with the people affected. In a small team, that means three audiences.

  • The owner of each affected system, as soon as containment is done, so nobody works on top of the damage.
  • The team, in one short message: what happened, what is stopped, what they should not touch, and when you will update them.
  • People outside, if data left: customers whose records were exposed, and anyone you have a contractual or legal duty to tell. Take advice on the legal part; it depends on your country and your contracts.

Write the update in plain words and name the agent as the actor. “Claude, acting for Sam, deleted six tasks at 14:10; all six are restored” is more useful to everyone than “a system issue”.

5. Learn: narrow what allowed it (ID.IM)

The cause of most agent incidents is access the agent did not need, or a check that was not there. Ask which single change would have stopped this one, make it, and write it down where the next session will read it.

  • Narrow the credential: fewer scopes, a lower role, an expiry date.
  • Add the missing control: a deny rule, a sandbox setting, a branch rule, an approval step. The settings are in security controls for AI coding agents.
  • Change the instructions the agent reads, so the rule is followed in every session rather than remembered by one person.
  • Add the incident to your next security review of AI agent access, so the change is checked.

The checklist

AI agent incident checklist
CONTAIN     stop the session; revoke the agent's tokens, not the person's
            remove keys and hooks it created; rotate secrets in its reach
ASSESS      save records: git, transcript, tool history, provider logs
            timeline: what it read, what it changed, where it could send
RESTORE     revert commits; restore data from a checked backup
            use each tool's undo, knowing what it does not cover
COMMUNICATE system owners, then team, then outside if data left
LEARN       one change that would have stopped it; write it down

The board side, on fenbs

For the board itself, fenbs covers most of the checklist. Containment is one click: revoking an assistant’s token in Settings, “Connect an AI assistant”, stops it at once, ends a sign-in’s refresh as well as its access, and leaves the person signed in. For assessment, History records every change with who made it, an assistant’s as “Claude via” the person it acted for, and can be filtered to AI assistants; what it did stays there after the token is revoked. For restoring, a deleted task is soft-deleted and fenbs_restore_item brings it back with its number, comments, followers and links. Files attached to it are removed when it is deleted, so they are the one thing to recover from elsewhere. For learning, the board’s Decisions page records what was decided, by whom and why, which is a good home for the rule you changed.

Related

What a useful record contains: an audit trail for AI agents. Limiting what an assistant can do on a board before anything goes wrong: how to keep an AI agent from wrecking your board and assistant tokens and scopes.

Questions people ask.

What is the first thing to do when an AI agent does something harmful?

Stop the session and revoke the agent’s own credentials in every system it could reach, leaving the person it worked for signed in if you can. Then rotate any secret it could have read. Investigate after that, not before.

Does NIST have guidance for AI agent incidents?

NIST SP 800-61 Revision 3, published in April 2025, is general incident response guidance rather than guidance specific to AI agents. It organises response around the Cybersecurity Framework 2.0 functions Detect, Respond and Recover, with preparation under Govern, Identify and Protect, and it applies to agent incidents as well as any other.

Can I undo everything a coding agent changed with rewind?

No. Claude Code’s checkpoints restore edits made by its own file editing tools, but not changes made by shell commands or by most subagents. Use version control to revert committed work and backups for data.

Should I rewrite git history to remove a secret an agent committed?

Rotate the secret first; that is what makes it harmless. Rewriting history is disruptive and may not remove copies others have already pulled, so decide on it separately and with care.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.