How to Reduce Claude Code Token Usage, and See What You Used

Where to see what a Claude Code session used, what actually spends the tokens, and the settings and habits that cut usage without making the work worse.

7 min read

To see what Claude Code used, run /usage for the session’s tokens, an estimated cost and your plan limits, and /context for what is filling the context window right now. To use fewer tokens, clear the conversation between unrelated tasks, match the model and effort level to the job, send verbose work to subagents on a smaller model, switch off MCP servers you are not using, and keep CLAUDE.md short. Just as important is what not to cut: the tests Claude runs to check its work, and planning on large changes, both cost tokens and save more than they spend.

This post is about usage and cost. How each mechanism decides what Claude sees, and what survives compaction, is covered in context engineering for Claude Code.

See what you used

  • /usage (also /cost and /stats) opens with a Session block: tokens by model, split into input, output, cache reads and cache writes, and an estimated cost. Anthropic’s cost documentation (opens in a new tab) says the figure is computed locally at list price and is meant for API users; on a subscription it is not what you are billed.
  • On Pro, Max, Team and Enterprise plans, the same screen shows your plan usage bars and a breakdown of recent usage by skill, subagent, plugin and MCP server. Press d or w to switch between the last day and the last week. It covers this machine only, not other devices or claude.ai.
  • The breakdown also flags any behaviour, such as long context or cache misses, that accounts for 10% or more of recent usage. That flag is usually the fastest answer to “why is this session so expensive?”
  • /context shows the window as a coloured grid, with suggestions for context-heavy tools, memory files that have grown too large, and warnings as you near capacity.
  • /insights writes a report on how you use Claude Code across recent sessions: what you work on and where sessions go wrong. The analysis itself counts against your usage.

To keep an eye on it without running a command, put the context percentage in the status line (opens in a new tab). Claude Code passes the script a JSON object that includes context_window.used_percentage:

~/.claude/statusline.sh (point statusLine.command at it in settings.json)
#!/bin/bash
input=$(cat)
pct=$(echo "$input" | jq -r '.context_window.used_percentage // 0')
echo "context $pct%"

For a team, per-person numbers come from outside the terminal: organisation analytics on Team and Enterprise plans, the Console on API billing, or an OpenTelemetry export on any setup.

What actually spends the tokens

Usage climbs in ways that are not obvious from what you typed. The costs page lists the usual reasons a long session draws more than expected:

  • The whole conversation goes with every request. Each tool call is another request carrying its results, so a one-line question at the end of a day-long session still pays to re-read the whole history, at the cached rate if the cache is warm.
  • Cache misses. The first message after a break longer than the prompt cache lifetime (opens in a new tab) reprocesses the full context. The lifetime is an hour on a subscription and five minutes by default on an API key.
  • File reads and command output. A test run that prints a thousand lines, or a search that opens forty files, lands in the window and stays there.
  • Thinking. Extended thinking is on by default and thinking tokens are billed as output tokens.
  • Subagents, workflows and agent teams. Each sends its own requests on top of the main conversation, and each teammate keeps spending until it exits.
  • Scheduled tasks. A /loop fires on its interval even while you are away, sending the full context each time.
  • Compaction. /compact has to read the conversation to summarise it, so compacting a large context is itself a large request. /clear costs nothing.

Settings that cut usage

  • Match the model to the job. Claude Code now starts on Opus on most plans and clouds; drop to Sonnet for routine edits and quick questions to spend less. Switch with /model, or use the opusplan alias, which runs Opus in plan mode and Sonnet for the edits. Claude Sonnet vs Opus covers which fits which job.
  • Lower the effort on simple work. /effort low or medium cuts thinking on routine tasks; raise it again when the problem is hard. The model configuration guide (opens in a new tab) lists the levels each model supports.
  • Put subagents on a smaller model. Add model: haiku to a subagent that only searches or summarises, or set CLAUDE_CODE_SUBAGENT_MODEL for subagents that name no model. The built-in Explore subagent inherits your session’s model; define your own Explore with model: haiku to keep searches cheap.
  • Switch off MCP servers you are not using. Tool definitions are deferred by default, so an idle server costs little, and /mcp disable <server> disconnects it altogether. Where a command-line tool such as gh does the job, it adds no tool listing at all.
  • Keep CLAUDE.md under about 200 lines. It loads in every session, so a workflow described there is paid for even when you are doing something else. Move procedures into skills, which load only when used; /skills can sort yours by token cost.
  • Filter noisy output before Claude sees it. A PreToolUse hook can rewrite a test command to print only failures, turning tens of thousands of tokens of output into a few hundred.
.claude/agents/log-reader.md
---
name: log-reader
description: Reads long logs and test output and reports only the failures that matter. Use for any output over a few hundred lines.
tools: Read, Grep, Glob
model: haiku
---
Read the file or output you are given. Report each distinct failure once,
with the file and line, and nothing else. Do not suggest fixes.

A subagent like that keeps the noise in its own window and returns a short summary to yours. More patterns are in Claude Code subagent examples.

Habits that cut usage

  • Run /clear when you switch to unrelated work. Use /rename first so you can /resume the old session if you need it.
  • Compact with a focus before a long new stretch, for example /compact keep the migration plan and the failing test names, so the summary keeps what you choose.
  • Ask side questions with /btw. The answer never enters the conversation, so it is not re-sent on every later message.
  • Be specific. “Add input validation to the login function in auth.ts” reads one file; “improve this codebase” reads dozens.
  • Stop a wrong turn early with Esc rather than letting Claude finish and then asking it to undo.
  • After a long break on Pro or Max, accept the offer to resume from a summary instead of the full history.
CLAUDE.md: tell compaction what to keep
# Compact instructions
When compacting, keep the list of modified files, the test commands,
and any error messages still unresolved.

What not to cut

  • Verification. Letting Claude run the tests costs tokens; a wrong fix you discover tomorrow costs a whole new session.
  • Plan mode on large or unfamiliar changes. A plan you correct before any edit is cheaper than an implementation you throw away.
  • The CLAUDE.md lines Claude cannot guess. Strip the build command and Claude spends more tokens finding it than the line ever cost.
  • Prompt caching. It can be switched off with an environment variable, but caching is what makes re-sending the conversation affordable. Leave it on.
  • History you still need. Mid-task, compact rather than clear; /clear is for when the old conversation is no longer useful.

Stop paying to re-explain the work

A quieter cost is the start of each session: Claude re-reading files, or you re-typing the background, to work out where things stand. Keep that state outside the session. With fenbs connected over MCP, one CLAUDE.md line, “at the start, call fenbs_get_context, then list the In Progress lane”, gives a new session the AI context notes earlier sessions left and the tasks already under way, in two tool calls. Keep those notes short and true; they load every time, just like CLAUDE.md. A task-tracking workflow for Claude Code shows the rest of the rhythm.

Related

The commands above, with the rest of the set: Claude Code commands cheat sheet. Why long sessions get worse as well as dearer: AI context rot. What fenbs keeps for assistants between sessions: AI context, connected through the Claude Code integration.

Questions people ask.

How do I check token usage in Claude Code?

Run /usage. It shows the current session’s tokens by model, including cache reads and writes, and an estimated cost. On subscription plans it also shows plan usage bars and a breakdown by skill, subagent and MCP server. /context shows what is in the context window right now.

Does /compact save tokens?

It makes later requests smaller, because they carry a summary instead of the full history, but the compaction itself is a large request because it reads the whole conversation. When you do not need the old conversation at all, /clear is cheaper.

Do MCP servers use tokens when I am not calling them?

A little. By default Claude Code lists only tool names and server instructions until a tool is needed, so an idle server costs far less than it once did. Disabling servers you are not using with /mcp removes even that.

Which model should I use to save tokens in Claude Code?

Sonnet for most coding tasks, with Opus kept for complex reasoning and architecture. For subagents that only search or summarise, Haiku is usually enough. Lowering the effort level on routine work also reduces thinking tokens.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.