MCP security risks: what can go wrong when an AI connects to your tools

Eight ways an AI assistant connected over MCP can be turned against the people it works for, each from a primary or reputable source, and the one mitigation that matters most for each.

Updated 8 min read

The main MCP security risks are prompt injection through content a tool returns; tool poisoning, where a tool’s description carries hidden instructions; rug pulls, where a server changes its tools after you approved them; over-broad tokens; stolen tokens; the confused deputy problem in servers that front other services; data exfiltration that crosses from one tool to another; and malicious or compromised community servers. None of these needs a flaw in the model. They follow from one fact: an assistant treats what an MCP server tells it as something to act on, and it holds real access to act with.

This piece is the list of what can go wrong. The companion piece, MCP security best practices for teams, is what to do about it. If you need the protocol explained first, see what MCP is.

Why MCP changes the picture

Before MCP, a language model mostly produced text, and a person decided what to do with it. With MCP, the model chooses tools and calls them. The specification (opens in a new tab) describes tools as model-controlled: the model discovers and invokes them based on its understanding of the conversation. OWASP’s Excessive Agency (LLM03 in the 2026 list, LLM06 in 2025) names the resulting class of problem: damaging actions performed in response to unexpected, ambiguous or manipulated output, made possible by too much functionality, too many permissions or too much autonomy. Each risk below is a specific way that happens.

1. Prompt injection through tool results

An assistant reads a web page, an email, a support ticket or an issue, and the text contains instructions. The model cannot reliably tell the difference between data it was asked to read and instructions it was given, so it may follow them. OWASP calls this indirect prompt injection (LLM01:2025 Prompt Injection), and its MCP Top 10 (opens in a new tab), still in beta, lists “Prompt Injection via Contextual Payloads” as MCP06:2025.

A documented case: in May 2025 Invariant Labs showed (opens in a new tab) that a malicious issue filed on a public GitHub repository could lead an agent using the GitHub MCP server to read the user’s private repositories and publish their contents in a pull request on the public one. They were explicit that this was not a bug in the server’s code but an architectural problem: one token could reach both public and private repositories.

Mitigation: limit what any single session can reach, so that content from an untrusted source cannot steer an assistant into data or actions that matter. Least-privilege tokens, one scope of work per session, and human approval for anything that publishes or sends.

2. Tool poisoning

Every MCP tool comes with a description, and the model reads it as guidance on when and how to use the tool. A malicious server can hide instructions there that the user never sees in the interface. Invariant Labs described this (opens in a new tab) in April 2025 with a harmless-looking “add” tool whose description told the model to read the user’s SSH key and MCP configuration file and pass them along as a parameter. OWASP lists it as MCP03:2025 Tool Poisoning, and the MCP specification says clients must treat tool annotations as untrusted unless they come from trusted servers.

Mitigation: connect only servers you trust, read the full tool descriptions before approving, and use a client that shows you the full inputs of a tool call before it runs.

3. Rug pulls: tools that change after approval

You review a server, its tools look fine, you approve it. Later the server changes its tool descriptions. The protocol supports this legitimately: a server can notify the client that its tool list has changed. Invariant Labs called the malicious version a rug pull and compared it to supply chain attacks on package registries. Your approval was of something that no longer exists.

Mitigation: for servers you run yourself, pin the version you reviewed and re-read the tools when you upgrade. For remote servers, prefer those run by the vendor of the underlying tool, whose own reputation is on the line.

4. Over-broad tokens

A token that can do everything turns any of the other risks into a large one. The specification’s security guidance (opens in a new tab) warns against wildcard or omnibus scopes and against requesting every scope up front, noting that a stolen broad token enables access unrelated to the job and is harder to revoke without disrupting everything. OWASP lists “Privilege Escalation via Scope Creep” as MCP02:2025.

Mitigation: the fewest scopes that do the job, a role no wider than the person needs, and a regular check that nothing has crept wider. A monthly routine for that check is in how to audit AI agents.

5. Token theft

An access token is a key. If it is pasted into a config file, committed to a repository, printed in a log or left on a lost laptop, whoever finds it can use it, and their requests look legitimate. The MCP authorization specification addresses this directly: clients and servers must store tokens securely, and authorization servers should issue short-lived access tokens to reduce the impact of a leak. OWASP’s first MCP risk, MCP01:2025, is “Token Mismanagement & Secret Exposure”.

Mitigation: sign in through the browser so the token is never copied by hand; keep hand-issued tokens in a secret store; and make sure every token can be revoked on its own, at once.

6. The confused deputy

This one is for teams running or choosing a server that sits in front of another service. The specification describes it in detail: an MCP proxy server uses one static client id with a third-party authorization server, lets MCP clients register dynamically, and the third party remembers consent in a cookie. An attacker registers a client with their own redirect address and sends the user a link. The consent screen is skipped because the cookie is there, and the authorization code goes to the attacker, who can then act as the user.

Mitigation: such proxy servers must get the user’s consent for each dynamically registered client before forwarding to the third party, match redirect addresses exactly, and never pass a client’s token through to the service behind them. As a buyer, ask whether a server you are considering does this.

7. Data exfiltration across tools

An assistant connected to several servers at once can be steered to move data between them: read it with one tool, send it with another. Invariant Labs showed a malicious server whose tool description changed how the agent used a separate, trusted email tool, redirecting messages to the attacker. The GitHub case in risk 1 is the same shape: read from private, write to public.

Mitigation: keep the tools that can read sensitive data and the tools that can send data outward from being available in the same session where you can, and require confirmation for anything that sends.

8. The supply chain of community servers

Many MCP servers are packages you download and run on your own machine, with your privileges. The specification notes that a local server can run arbitrary code with the client’s privileges and that users may have no visibility of what it executes. OWASP lists “Software Supply Chain Attacks & Dependency Tampering” as MCP04:2025.

A documented case: in September 2025 an npm package called postmark-mcp, impersonating the email company Postmark, was found to blind-copy every email it sent to an outside address. Postmark confirmed (opens in a new tab) it was not theirs and that the package had built trust over fifteen releases before the backdoor appeared in version 1.0.16. Anthropic’s own documentation notes that it does not security-audit or manage any MCP server.

Mitigation: install servers from the vendor of the tool, pin versions, read what changed before upgrading, and keep a written list of approved servers.

The risk that makes the others worse: no record

Every risk above is easier to contain if you can see what happened. OWASP lists “Lack of Audit and Telemetry” as MCP08:2025. If an assistant’s changes are recorded as yours, or not recorded at all, you cannot tell a manipulated assistant from a mistaken colleague. The case for a record kept by the tool itself is in an audit trail for AI agents.

Where a task board fits in this list

A task board is content your assistants read, so it is subject to risk 1 like any other source: a task written by someone outside the team could contain instructions. On fenbs that is bounded in three ways. Who can add or edit tasks is decided by role, so a client can be given a role that only reads and comments. An assistant’s token is capped by both its scopes and its owner’s role, so injected text cannot make it do what its owner could not. And every change is recorded as “Claude via” the person it acted for, with task deletes being soft and restorable. None of that stops an assistant from reading a hostile sentence. It limits what the sentence can achieve and makes the result visible.

Questions people ask.

What is the biggest security risk with MCP?

Prompt injection is the most widely discussed, because any content an assistant reads can carry instructions. It becomes serious when combined with a token that can reach sensitive data and a tool that can send data out, so limiting scopes and requiring approval for sends reduces most of the harm.

What is MCP tool poisoning?

It is when a server hides instructions in the description of one of its tools. The model reads descriptions as guidance, so it may follow them, even though the user never sees them. OWASP lists it as MCP03:2025.

What is an MCP rug pull?

It is when a server changes its tool descriptions after you have reviewed and approved it, so what you approved is no longer what runs. Pinning versions of local servers and re-reading tools on upgrade are the usual defences.

Are remote MCP servers safer than local ones?

They carry different risks. A local server runs code on your machine with your privileges, so a malicious package can do a lot of damage. A remote server cannot touch your machine but holds a token to your data, so how it signs you in, what it lets you narrow and how it records changes matter more.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.