Indirect Prompt Injection: How It Reaches AI Agents

An AI agent does not need to be talked into anything by its user. It only needs to read something an attacker wrote. Where that text comes from, how an attack unfolds, why a better prompt will not stop it, and the controls that do limit the damage.

8 min read

Indirect prompt injection is an attack in which instructions are hidden in content an AI system reads, rather than typed by the person using it. The attacker never talks to the model. They plant text in a web page, an email, a document, an issue or a tool’s output, wait for an agent to read it as part of a legitimate job, and rely on the model treating some of that text as instructions. For an assistant that only answers questions, the result is a wrong answer. For an agent that can call tools, it can be a sent email, a changed record or data copied somewhere it should not go.

Where the term comes from

The name was set out in a February 2023 paper by Kai Greshake and colleagues, Not what you’ve signed up for (opens in a new tab). Their central observation was that applications built on language models blur the line between data and instructions, so an adversary can exploit them remotely, without any direct interface, by placing prompts in data the application is likely to retrieve. They demonstrated it against real systems, including Bing’s GPT-4 chat, and listed consequences that now read like a description of agents: data theft, self-spreading attacks, and manipulation of which APIs the application calls.

OWASP puts both kinds under a single entry, LLM01:2025 Prompt Injection (opens in a new tab). Direct injection is when the user’s own input changes the model’s behaviour. Indirect injection is when the model accepts input from external sources, such as websites or files, and something in that content changes its behaviour. OWASP adds two points worth keeping: the injected text does not need to be visible to a person, only readable by the model, and retrieval or fine-tuning does not fully remove the risk.

How it reaches an agent

Any text an agent reads on someone else’s behalf is a channel. The common ones:

  • Web pages. A research or browsing agent summarises a page that contains text meant for the model, sometimes styled so a human visitor never sees it.
  • Emails and calendar invites. An assistant that triages an inbox reads every message, including ones written by strangers.
  • Documents and spreadsheets. A shared file, a CV, a supplier PDF or a comment in a spreadsheet cell is read in full, whoever wrote it.
  • Issues, pull requests and code comments. A coding agent asked to “fix the open issues” reads text that anyone with an account may have written.
  • Tool results. Whatever an MCP server or API returns becomes part of the model’s context. That route, along with poisoned tool descriptions, is covered in MCP security risks.
  • Memory. An injection that gets written into notes the agent reads next time outlives the session that saw it. That is its own subject: agent memory security.

How an attack unfolds, step by step

The payload itself matters less than the shape of the path. Nearly every documented case follows the same four steps, and each step is a place to break the chain.

  1. Placement. The attacker puts text somewhere the agent will look: a public issue, a web page on a topic the agent researches, an email to an address the assistant reads.
  2. Retrieval. A person gives the agent an ordinary job, and the job leads it to that content. Nobody pastes the attack in; the agent fetches it.
  3. Hijack. The model reads the planted text in the same context as its real instructions and gives it weight. It may be framed as a note from the user, a system message or an urgent correction.
  4. Action. The agent uses tools it legitimately holds: it reads a private file, calls a send or publish tool, edits a record, or puts data into a URL it fetches.

Three worked paths show the pattern without needing any exploit text. A coding agent is asked to triage public issues; one issue tells it to gather details from the owner’s private repositories and include them in a pull request, and it has a token that can read both. An inbox assistant summarises the morning’s mail; one message asks it to forward the latest invoice thread to an outside address, and it holds a send tool. A research agent reads a product page whose hidden text tells it to recommend that product and describe competitors poorly; nothing is stolen, but the report the person reads has been written by someone else.

In each case the damage comes from a combination: untrusted content, access to something valuable, and a way to act or send. Remove any one and the same text becomes harmless. That combination is the thread running through AI agent security best practices.

Why prompt defences alone fail

The obvious fix is to tell the model not to follow instructions it finds in content: wrap retrieved text in markers, add a system message, fine-tune it to be suspicious. These help at the margin and are worth doing, but they cannot be the boundary, for three reasons.

  • The model has one input. Instructions and data arrive as the same kind of text in the same context. A delimiter is a convention the model is asked to respect, not a wall it cannot cross.
  • Attackers adapt. A 2025 study by researchers from several labs, The Attacker Moves Second (opens in a new tab), bypassed twelve recently published defences with attack success rates above 90% for most of them, although most had originally reported near-zero rates. The defences had been tested against fixed attack strings rather than an attacker tuning against them.
  • Low rates still add up. Anthropic reports steady progress with training, classifiers and red-teaming, and still describes prompt injection as far from a solved problem (opens in a new tab), adding that no browser agent is immune. An agent that reads thousands of pages meets the rare case eventually.

The practical conclusion is to assume that some injected text will be followed, and to design so that following it does little harm.

The controls that help

The research direction agrees with that conclusion. A June 2025 paper by authors from several companies and universities, Design Patterns for Securing LLM Agents against Prompt Injections (opens in a new tab), states the principle plainly: once an agent has ingested untrusted input, it must be constrained so that the input cannot trigger consequential actions. The controls below are ways of doing that with ordinary tools.

Least privilege, per job

Give each agent the narrowest credential for the job in front of it. An agent that triages public issues does not need a token for private repositories. An agent that summarises email does not need a send tool. The less it holds, the less an injection can use. How to set that up on a board, with scopes and roles, is in how to keep an AI agent from wrecking your board.

Human confirmation for consequential actions

Anything that sends, publishes, pays, deletes or changes access waits for a person, and the person sees the full action, not the agent’s summary of it. This is the one control injected text cannot talk its way past, provided the approval screen shows what will actually happen. When to ask and when not to is covered in human in the loop for AI agents.

Separate data from instructions in the design

Rather than asking one model to read untrusted content and decide what to do, split the two. Decide the plan before reading anything untrusted, then let the untrusted content fill in values but not choose actions. CaMeL (opens in a new tab), a 2025 design from researchers at Google and ETH Zurich, takes this furthest, separating control flow from data flow so that retrieved data can never change which tools are called; on the AgentDojo benchmark it completed 77% of tasks with provable security, against 84% for the undefended agent. For a small team the lighter version is structural: a reading step with no write tools, and a separate acting step that only sees what the reading step extracted.

Egress limits

Many attacks end in sending data somewhere: an email, a web request, a public comment, an image URL with data in its query string. Limit network access to an allowlist of domains, keep send and publish tools off in sessions that read untrusted content, and treat an unexpected outbound request as an incident signal.

Provenance

Know where each piece of content came from and who wrote it, and record which agent did what. When something odd happens, the first question is what the agent read just before it; a signed record of its actions turns that from guesswork into a timeline. The response steps are in AI agent incident response.

Where a task board fits

A task board is content your assistants read, so a task or comment written by the wrong person is a channel like any other. On fenbs the damage is bounded by the controls above. Who can create or edit tasks is set by role, so a client can be given a role that only reads and comments. An assistant’s token is capped by both its scopes (read, write, comment) and its owner’s role, and moving tasks between lanes is a permission of its own, so an injected “mark all of these Completed” fails for an assistant that was never given it. Every change is recorded in History as “Claude via” the person it acts for, and revoking the token stops it at once without signing that person out. On a task pre-approved for AI, comments added after approval are information, not instructions.

Related

The MCP-specific risks: MCP security risks. The OWASP lists in plain words: OWASP guidance on AI agents. What each connection may do: assistant tokens and scopes. Testing an agent against injected content is part of AI agent evaluation.

Questions people ask.

What is the difference between direct and indirect prompt injection?

In direct prompt injection the person typing into the AI system is the one trying to change its behaviour. In indirect prompt injection the instructions are hidden in content the system reads for someone else, such as a web page, email, document or tool result, so the attacker never interacts with the model directly.

How does an indirect prompt injection attack occur?

An attacker places text where an agent will read it, a person gives the agent an ordinary job that leads it to that content, the model treats part of the text as instructions, and the agent then uses tools it legitimately holds to act on them, for example by reading private data or sending a message.

Can a better system prompt stop indirect prompt injection?

It can reduce it but not stop it. Instructions and data reach the model as the same kind of text, and published defences have been bypassed by attackers who adapt to them. Treat prompt hardening as one layer and rely on limits to what the agent can do.

What is the single most effective control?

Do not let an agent that reads untrusted content take consequential actions on its own. Give it the narrowest access for the job and require a person to confirm anything that sends, publishes, pays, deletes or changes access.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.