How to audit AI agents in a small team
A record of what your AI agents did is only useful if somebody reads it against what they were meant to do. Here is a monthly audit a team of two to twenty can run in under an hour.
8 min read
To audit AI agents in a small team, run the same seven checks once a month: list every assistant that is connected and who connected it; check what each one is allowed to do; revoke the ones nobody uses; read what each one did against what it was for; sample its finished work for quality; check what it has written down for the next assistant; and record what you found and changed. For a team of two to twenty people that is under an hour, and the output is a short note that says who had access to what, and whether they used it well.
An audit is not the same as an audit trail
An audit trail is a record: every change, signed by whoever made it, kept by the tool rather than by the agent. We covered what a good one looks like in an audit trail for AI agents. An audit is the habit of reading that record on a schedule, against a question. The trail can be perfect and still useless if nobody opens it; the audit is what makes somebody open it.
The question is always the same: does what each assistant can do, and what it actually did, still match what we meant when we connected it? Access drifts. People connect an assistant for one job, the job ends, and the token stays. Someone widens a scope for a busy week and forgets to narrow it. None of this is dramatic, and all of it is found by looking.
Before you start
- One owner. In a small team the audit belongs to one person, usually whoever runs the board. Rotating it sounds fair and means nobody remembers how it was done last time.
- A fixed date. The first working Monday of the month, say. An audit that happens “when things are quiet” does not happen.
- A place for the result. A task on the board works well: the findings go in the note, the follow-ups become tasks of their own, and next month’s audit starts by reading last month’s.
The checklist
- Inventory: which assistants are connected, and by whom.
- Permissions: what each connection can do.
- Stale access: what has not been used, and what belongs to someone who has left.
- Activity against intent: what each assistant actually did.
- Quality: a sample of its finished work, checked properly.
- Shared context: what the assistants have written down for each other.
- Record and act: revoke, narrow, rotate, and write it down.
1. Inventory
Ask each person to list the assistants they have connected to the team’s tools, then compare that with what the tools themselves report. The two lists rarely match on the first run. There will be a connection from a trial nobody finished, or a script someone set up on a laptop. The OWASP MCP Top 10 (opens in a new tab), a project still in beta, calls unapproved servers and connections outside a team’s oversight “Shadow MCP Servers” (MCP09:2025). In a small team the fix is not a scanner; it is asking.
Check the client side too. In Claude Code each person can print what their machine is connected to:
claude mcp list # every MCP server this machine has configured claude mcp get fenbs # details for one of them
Servers configured for a whole project (opens in a new tab) live in a .mcp.json file in the repository, so they show up in version control and can be reviewed like any other change.
2. Permissions
For each connection, write down what it can actually do. On most tools that is two things multiplied together: the role of the person it acts for, and the scopes the token was given. An assistant approved with write access by an Owner can do a great deal more than one approved with read access by a Viewer, even if both are “Claude”.
Then ask whether each one still needs what it has. The usual findings are a write scope on an assistant that only ever reads and reports, and an assistant acting for someone whose own role is wider than their work needs. Narrow the role or reissue the token with fewer scopes. The reasoning behind this is in roles and permissions for humans and AI agents; the audit is where you check it is still true.
3. Stale access
Look for three things: connections that have never been used, connections not used for a month or more, and connections belonging to people who have left or changed role. Revoke the first and third without discussion. Ask about the second, and revoke unless there is a reason. A token that is not being used is pure risk: it grants access and does no work.
4. Activity against intent
Now read what each assistant did this month, one assistant at a time, next to a one-line statement of what it is for. “Cursor on Priya’s account: picks up bugs, fixes them, comments with the commit.” Then look for anything outside that line. The things worth stopping on are deletions, moves to Completed, changes to who is on the board, and bulk activity: forty changes in ten minutes is either a good import or a loop, and you want to know which.
Also ask each person whether their assistant reported being refused anything. A refusal is either a role that is too narrow for a reasonable job, or an assistant attempting something it should not. Both are findings.
5. Quality
Pick five tasks each assistant moved to Completed and check them properly. Does the comment say what changed, name the commit, and say how it was verified? Does the commit exist? Was the plan written before the work started? Then pick five tasks it created: are they real, are they the right kind (feature, enhancement or bug), and are any of them duplicates? Five is enough to spot a pattern and small enough that you will actually do it.
6. Shared context
If your assistants read shared notes before they start, those notes steer everything they do. Read the ones added or changed this month. Is each still true? Does any contradict another? Every assistant that reads a note repeats it faithfully, so a wrong note is worth more to fix than anything else on this list.
7. Record and act
Write the result down in the same shape every month, so months can be compared. Then act on it the same day: revoke what is stale, reissue what is too wide, and file a task for anything that needs more than a minute.
AI assistant audit, September Connected: 6 (Claude Code x3, Cursor x2, nightly triage script) Revoked: 2 (one never used; one belonged to a contractor who left) Narrowed: 1 (Cursor on Priya: read, write, comment -> read, comment) Sampled: 10 completed tasks; 9 named a commit and a test run Context: 1 note out of date, corrected Follow-ups: ENH-041, BUG-042
Rotation belongs here too. A token issued by hand for a script should be revoked and reissued on a schedule you choose, and at once if the machine it lives on is lost or the person who set it up leaves. A connection made by signing in is rotated by revoking it and signing in again from the assistant.
How fenbs supports each step
On a fenbs board the raw material for the audit is already in two places.
- Settings, “Connect an AI assistant”, lists every token you have issued or approved by name, with its scopes (read, write, comment) and when it was last used, or “never used”. Revoked tokens stay on the list, struck through with the date, so the inventory includes what used to have access. Revoke is one click; it stops the token at once and leaves your own sign-in alone.
- The Your connections page, under your account, lists what is connected now and every board an assistant acting as you could reach, with your role on each. That is steps 1 and 2 on one screen, per person.
- History can be filtered to AI Assistants only, and searched by a ref, a name, a project or a verb. Every entry reads “Claude via” the person it acted for. Changes to who is on the board are recorded there too, and a deleted task can be restored from its line.
- A token’s power is the narrower of its scopes and its owner’s role, checked on every call. Narrow someone’s role during the audit and every token they issued narrows with it immediately.
- AI context notes are signed by whoever wrote them, person or assistant, so step 6 is a matter of reading the notes with an assistant’s name on them.
The full trail of who changed what is part of the Teams plan; see pricing. Tokens are per person, so on a team board each person runs step 1 and 2 for their own connections and the owner collects the results.
What the audit does not cover
A board audit covers the board. An assistant that also works in your code, your email or your cloud account needs the same questions asked of those systems, using their own records: version control history for code, the provider’s access logs for the rest. Quoting a task’s ref in each commit message is what lets you walk from a line in History to the change it describes. And if something in the audit looks like a security problem rather than housekeeping, stop auditing and revoke first; the record will still be there when you come back to it.
Related
The controls this audit checks are set up in how to give an AI agent access to your project board. For the wider picture of what to watch for when assistants connect over MCP, see MCP security best practices for teams. Scopes are defined in assistant tokens and scopes.