What AI Agents Really Cost to Run (Without Price Lists)

An AI agent’s bill is mostly tokens, and it grows with the number of steps and the size of the context at each step, not with the length of your request. What drives the cost, how caching and batching cut it, the limits that stop a runaway agent, and how to track cost per task.

8 min read

An AI agent costs what its tokens cost, plus whatever its tools charge, plus the time people spend checking its work. The tokens are the part that surprises people. An agent works in steps, and at every step the model is sent the whole working context again: its instructions, its tool definitions, the conversation so far and every tool result it has collected. So the cost of a task grows with the number of steps and with how large the context has become by each one, not with how short your request was. Anthropic’s engineering team reported that agents typically use about four times the tokens of a chat, and multi-agent systems about fifteen times. The levers, in order of payoff: fewer steps, less in the context, caching, a smaller model for simple sub-work, batching what can wait, hard limits, and tracking cost per task so you know which of these to pull.

There are no prices here, because they change often and differ by model and plan; check your provider’s own pricing page. MCP’s share of the bill, from tool definitions to oversized results, is covered in MCP token usage, and Claude Code’s own commands and settings in how to reduce Claude Code token usage.

What you pay for

  • Input tokens: everything sent to the model on each request. For an agent this is the large number, because it is re-sent at every step.
  • Output tokens: what the model writes, including tool calls. On current models output is priced well above input per token.
  • Thinking tokens: reasoning the model does before answering. Anthropic’s extended thinking documentation reports them as part of the billed output tokens.
  • Cache writes and reads: storing a prompt prefix costs a little more than normal input, and reading it back costs much less.
  • Tool fees: some hosted tools, such as web search or code execution, are billed separately from tokens.
  • Everything around the model: the APIs your tools call, sandboxes or servers the agent runs on, and the hours people spend reviewing its output.

On a subscription, such as a Claude or ChatGPT plan used through Claude Code or Codex, you do not see a per-token bill; you see usage limits that reset on a schedule. The same drivers decide how fast you reach them.

Why one task costs ten times another

As an illustration of the arithmetic, not a measurement: an agent starts with 10,000 tokens of instructions and tool definitions, and each step adds about 2,000 tokens of tool results. By the twentieth step it sends about 48,000 tokens, and the twenty steps together send about 580,000 tokens of input for a task whose request was one sentence. Double the steps to forty and the total more than triples. That is why the drivers below matter more than the length of your prompt.

  • Steps. Every tool call is another round trip carrying the whole context. A vague task that sends the agent exploring costs more than a precise one that tells it where to look.
  • Context growth. Large tool results, file reads and command output stay in the conversation and are paid for again on every later step.
  • Retries and wrong turns. A failing test run three times, or a tool called with the wrong arguments and retried, is paid in full each time.
  • Parallel agents. Each subagent or teammate has its own context and re-reads it every step. In its write-up of its multi-agent research system (opens in a new tab), Anthropic says token usage alone explained 80% of the variance in performance in its evaluation, which is why more agents can be worth it for hard research and wasteful for routine work.
  • Model choice. Larger models cost more per token. A small model is often enough for searching, summarizing and classifying.
  • Idle gaps. A cached prefix expires after a few minutes by default, and the first request after that pays full price for the whole context again.

Caching: the biggest single lever

Because an agent re-sends the same prefix at every step, caching that prefix is the largest saving available without changing what the agent does. Anthropic’s prompt caching documentation (opens in a new tab) prices cache reads at a tenth of the normal input rate on most models, and lower on some newer ones. Writing to the cache costs a quarter more than normal input for the default five-minute lifetime and double for a one-hour lifetime, and each use refreshes the lifetime at no extra cost.

  • Order matters. The cache follows tools, then system prompt, then messages. Change a tool definition and every cached level after it is invalidated; change the system prompt and the message cache goes too.
  • Keep the front stable. Put fixed instructions and reference material first, and anything that changes, such as a timestamp or the user’s name, at the end.
  • Watch the hit rate. A low cache-read count usually means something near the top of the prompt changes on every request.

OpenAI turns prompt caching on by default for supported models. Its prompt caching guide (opens in a new tab) gives the same advice to put stable instructions and shared material first, and reports cached tokens in usage.input_tokens_details.cached_tokens, so you can see whether caching is working.

Batch what can wait

Not every agent job needs an answer now. Nightly classification, bulk summaries and evaluation runs can go through a batch interface instead. Anthropic’s Message Batches API (opens in a new tab) processes requests asynchronously at half the standard cost, with most batches finishing in under an hour. It suits single requests better than long interactive loops, so it fits the steps of a pipeline more than a coding session.

Limits that stop a runaway agent

Cost control is not only about making each step cheaper. An agent stuck in a loop can spend a day’s budget in an hour, so put ceilings in three places:

  • Per task. Agent frameworks let you cap the number of turns. The Claude Agent SDK has maxTurns and a maxBudgetUsd option that stops a query when its estimated cost reaches the value you set; the OpenAI Agents SDK has max_turns.
  • Per key or workspace. The Claude Console lets you set your own monthly spend limit below your tier’s cap, and separate spend and rate limits for each workspace, so one agent’s workspace cannot use up the whole organization’s allowance.
  • Per person. Someone owns each agent and reads its cost every week. A limit nobody watches only tells you afterward.
Claude Agent SDK (TypeScript): cap turns and estimated spend for one task
import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const message of query({
  prompt: "Summarize the open issues labeled billing",
  options: { maxTurns: 20, maxBudgetUsd: Number(process.env.TASK_BUDGET_USD) },
})) {
  if (message.type === "result") {
    console.log(message.subtype, message.total_cost_usd, message.usage);
  }
}

Anthropic’s guide to tracking cost in the Agent SDK (opens in a new tab) is explicit that total_cost_usd is a client-side estimate from a bundled price table, fine for budgeting and development but not billing data. When a query stops for its budget, the result’s subtype is error_max_budget_usd.

Track cost per task, not per month

A monthly invoice tells you that agents cost something. It does not tell you which job, which agent or which prompt change caused it. Give each agent its own API key or workspace, and record tokens or estimated cost against each task it finishes. Anthropic’s Usage and Cost API (opens in a new tab) returns usage broken down by model, workspace and API key, separating uncached input, cached input, cache writes and output, with costs by workspace. Those splits answer the useful questions: which agent is expensive, whether caching is working, and whether last week’s change made things better.

  • Cost per completed task, not per request. A cheap agent that fails half the time is expensive.
  • Tokens per step, to spot context that keeps growing.
  • Cache hit rate, to catch a prompt prefix that changes on every call.
  • Review time, because an agent that saves tokens but produces work a person must redo has only moved the cost.

For a wider view of what to capture from a running agent, see AI agent observability.

The cheapest tokens are the ones you do not send

Two quiet costs come from coordination rather than from any single run. The first is re-explaining: each new session reads files, or is told the background again, to work out where things stand. The second is duplicate work: two agents, or an agent and a person, doing the same task because neither could see the other had started.

A shared board addresses both. With fenbs connected over MCP, a session can start with fenbs_get_context, which returns the rules and the AI context notes earlier sessions left, then list the In Progress lane to see what is already under way. A task holds the problem in its note, the steps in its plan, and a test status with test notes when it is done, so the next session reads a short record instead of rediscovering the work. fenbs does not measure tokens or cost; put the cost of a run in the task’s test notes if you want it next to the result. Keep context notes short, since they load every time. Cost-cutting work, such as “trim the triage agent’s tool results”, is an enhancement like any other, with a priority from 1 to 10.

Related

Where MCP spends tokens: MCP token usage. Claude Code settings that cut usage: how to reduce Claude Code token usage. Why several agents cost more: can AI agents talk to each other. What fenbs keeps between sessions: AI context.

Questions people ask.

How much does it cost to run an AI agent?

It depends mostly on how many steps a task takes and how large the context is at each step, because the model is sent the whole context again every time. Two tasks with the same one-line request can differ in cost many times over. Measure cost per completed task on your own workload rather than relying on a general figure, and check your provider’s current pricing.

Why do AI agents use so many more tokens than chat?

A chat is usually one request and one answer. An agent loops: it calls a tool, reads the result, and sends everything back to the model before the next step, so the same instructions and history are paid for many times. Anthropic has reported that agents typically use about four times the tokens of chat, and multi-agent systems about fifteen times.

What is the fastest way to cut an AI agent’s costs?

Make sure prompt caching is working, since an agent re-sends the same prefix at every step and cached reads cost a fraction of normal input. Then cut steps with more precise tasks, trim large tool results, use a smaller model for simple sub-work, and batch jobs that do not need an immediate answer.

How do I stop an AI agent from overspending?

Set limits at three levels: a cap on turns or estimated spend for each task, a spend limit for the API key or workspace the agent uses, and a named person who reviews its cost every week. Framework budget options are estimates, so rely on the provider’s own limits for a hard stop.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.