Security Controls for AI Coding Agents

A coding agent runs commands with your shell, your files and your network. Seven controls, each enforced by software rather than by a polite instruction, keep what it can reach small and what it did visible.

8 min read

The AI coding agent security controls that do the most work are seven, and each is enforced by the tool rather than by the model: permission rules that deny the few things the agent must never do; an operating-system sandbox around every command it runs; an allowlist of network destinations; secrets kept where neither the agent nor its commands can read them; branch rules that stop agent code reaching your main branch without a person’s approval; narrow tokens for every MCP server it connects to; and logs you can read afterwards. Instructions in a rules file shape what the agent tries. These controls decide what it can actually do.

This piece is about coding agents in particular, such as Claude Code in a terminal or GitHub Copilot’s cloud agent, and about the settings behind each control. The broader rules for every kind of agent are in AI agent security best practices, which also explains Claude Code’s permission modes; they are not repeated here.

Why coding agents need controls of their own

A chat assistant produces text and a person decides what to do with it. A coding agent reads your repository, runs shell commands, installs packages and pushes branches, all with whatever your machine is signed in to. Three things follow. Anything it reads, including a README or an issue written by a stranger, can carry instructions. Any command it runs can start other programs that the agent’s own permission checks never see. And anything it can reach on the network is somewhere your code could be sent. The controls below answer those three facts one layer at a time.

1. Deny rules for the things that must never happen

Start with a short list of actions no session should take without a person, and write them as rules rather than as advice. In Claude Code, permission rules (opens in a new tab) are evaluated deny first, and a deny set at any level cannot be allowed back by another level. An administrator can also switch off the modes that skip prompts, with permissions.disableBypassPermissionsMode set to "disable" in managed settings, where nobody can override it.

.claude/settings.json
{
  "permissions": {
    "deny": ["Read(./.env)", "Read(./.env.*)", "Bash(git push --force *)"],
    "ask": ["Bash(git push *)", "Bash(npm publish *)"],
    "disableBypassPermissionsMode": "disable"
  }
}

Know the limit before you rely on it. The same documentation calls Bash patterns that try to constrain arguments fragile: a rule for curl does not match the same program called by another path or inside sh -c. Treat command rules as a speed bump for honest mistakes, and put the real boundary in the sandbox.

2. A sandbox around every command

A sandbox is enforced by the operating system for a command and every process it starts, so it holds even when a rule is sidestepped or the agent has been talked into something. Claude Code’s Bash sandbox (opens in a new tab) runs on macOS, Linux and WSL2 and is switched on with /sandbox. Two settings turn it from a convenience into a control.

  • allowUnsandboxedCommands: false. By default, when a command fails inside the sandbox, the agent may retry it outside, through the normal permission prompt. This setting removes that escape hatch, and the panel calls it strict sandbox mode.
  • failIfUnavailable: true. By default, if the sandbox cannot start, commands run unsandboxed with a warning. For a managed rollout, make that a hard failure instead.

Remember what the sandbox covers. It isolates shell commands and their children; Claude’s own file tools are governed by permission rules instead. You need both, which is the point of layering them.

3. An allowlist for network egress

Egress is where a leak becomes a loss, so limit where commands can connect. Claude Code’s sandbox routes network traffic through a proxy that checks each host against allowedDomains; with strictAllowlist it refuses anything else instead of asking, and allowManagedDomainsOnly in managed settings stops developers widening the list. The documentation adds two warnings worth copying into your own notes: allowing a broad domain such as github.com can itself be a path for exfiltration, and because the proxy does not inspect encrypted traffic by default, techniques such as domain fronting may reach hosts outside the list.

Cloud agents have the same control in a different place. GitHub’s firewall for Copilot cloud agent (opens in a new tab) limits internet access with a recommended allowlist of package registries, and organisation owners and repository administrators can add to it. Read its stated limits: it applies to processes the agent starts through its Bash tool, not directly to MCP server processes or to processes started in the setup steps, and GitHub warns that disabling it lets the agent connect to any host.

4. Secrets out of reach, not just out of the prompt

Keeping secrets out of prompts is the start. The harder part is that commands inherit your environment and can read your home directory. By default Claude Code’s sandbox still allows reads of files such as ~/.ssh/ and ~/.aws/credentials, and sandboxed commands inherit the parent environment. The credentials block fixes both: deny entries block the files and unset the variables before each command runs, and deny entries from any settings scope add up, so no scope can remove one.

settings.json: credentials kept from commands
{
  "sandbox": {
    "enabled": true,
    "allowUnsandboxedCommands": false,
    "network": { "allowedDomains": ["registry.npmjs.org"] },
    "credentials": {
      "files": [
        { "path": "~/.ssh", "mode": "deny" },
        { "path": "~/.aws/credentials", "mode": "deny" }
      ],
      "envVars": [{ "name": "GITHUB_TOKEN", "mode": "deny" }]
    }
  }
}

If you run agents in a dev container, the same principle applies to what you mount. Anthropic’s guidance for containers is to avoid mounting host secrets such as ~/.ssh or cloud credential files, and to prefer repository-scoped or short-lived tokens, because a session running without prompts can read anything inside the container.

5. Branch rules and a person’s approval

Every control above can fail, so the last one sits where the code lands. Agent work goes on a branch and reaches the main branch only through a reviewed pull request. On GitHub, a ruleset (opens in a new tab) can require a pull request before merging, with required approvals and approval from someone other than the last pusher, and can block force pushes, restrict branch deletions and require status checks to pass.

GitHub builds several of these into its cloud agent. According to its risks and mitigations page (opens in a new tab), only people with write access can start it, it pushes to a single branch, it cannot mark its pull requests ready for review, approve them or merge them, and the person who asked for the work cannot approve the result. Workflows do not run on its code until someone with write access approves them. For a local agent, the ask rule on git push in the example above gives you the same pause.

6. Narrow tokens for every MCP server

Each MCP server a coding agent connects to is another system it can act in, with whatever that server’s token allows. Decide which servers are allowed at all (Claude Code’s managed settings accept allowedMcpServers and deniedMcpServers), then give each connection the narrowest scopes the job needs. The fuller list of practices is in MCP security best practices for teams.

A task board is a good example of what narrow looks like. On fenbs, a sign-in from an assistant gives it an access token that lasts an hour and is renewed with a refresh token, so nothing long-lived sits in a config file. A token issued by hand in Settings for a script has a name, the scopes read, write and comment, and an optional expiry of 7, 30, 90 or 365 days. Either kind is capped by the role of the person who connected it, and revoking it in Settings ends it, refresh and all, without signing that person out.

7. Logs you can read afterwards

Controls without records leave you guessing after an incident. Two kinds of log matter. The agent’s own: Claude Code can export OpenTelemetry events (opens in a new tab), including tool decisions and tool results, once CLAUDE_CODE_ENABLE_TELEMETRY is set; tool parameters are redacted unless you opt in, and an administrator can lock the collector endpoint in managed settings. A PreToolUse hook can also block a call outright, and allowManagedHooksOnly keeps developers from adding hooks of their own.

And the systems’ own: version control for code, and for every tool the agent writes to, a history kept by that tool under the agent’s name. On fenbs, History records each change as “Claude via” the person it acted for, and can be filtered to AI assistants. Why that record must come from the tool and not the agent is in an audit trail for AI agents.

The seven on one page

Coding agent controls
1 Deny rules     secrets, force push; ask on push and publish; bypass off
2 Sandbox        on, strict mode, fail if unavailable
3 Egress         allowlist only; no broad domains; cloud agent firewall on
4 Secrets        credential files and env vars denied; nothing mounted
5 Branches       PR required, approval by someone else, no force push
6 MCP            allowed servers only; narrowest scopes; expiry set
7 Logs           telemetry to a collector you control; tool-side history

Related

Connect a coding agent to a board with the smallest scopes using the Claude Code or GitHub Copilot steps. Scopes and roles are defined in assistant tokens and scopes, and checking these controls on a schedule is covered in how to audit AI agents.

Questions people ask.

Are permission rules enough to secure a coding agent?

No. Claude Code’s own documentation calls Bash argument patterns fragile, because the same program can be called by another path or inside another shell. Use rules for the obvious cases and put the real boundary in an operating-system sandbox with a network allowlist.

Does a sandbox stop a coding agent reading my SSH keys?

Not by default. Claude Code’s sandbox allows reads of most of the machine, including ~/.ssh and ~/.aws/credentials, unless you add them to the credentials or denyRead settings. Add them explicitly and unset secret environment variables for sandboxed commands.

Can a cloud coding agent merge its own pull request?

GitHub Copilot cloud agent cannot. GitHub documents that it cannot mark its pull requests ready for review, approve or merge them, and that the person who asked for the work cannot approve it either. Add a ruleset requiring approval so the same holds for every other route to your main branch.

Which network domains should a coding agent be allowed to reach?

Only the package registries and services the work needs, listed by name. Avoid broad entries such as all of github.com, which the Claude Code documentation notes can become a path for exfiltration, and prefer settings that refuse unknown hosts rather than asking.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.