How to Reduce Claude Code Token Usage, and See What You Used
Where to see what a Claude Code session used, what actually spends the tokens, and the settings and habits that cut usage without making the work worse.
7 min read
To see what Claude Code used, run /usage for the session’s tokens, an estimated cost and your plan limits, and /context for what is filling the context window right now. To use fewer tokens, clear the conversation between unrelated tasks, match the model and effort level to the job, send verbose work to subagents on a smaller model, switch off MCP servers you are not using, and keep CLAUDE.md short. Just as important is what not to cut: the tests Claude runs to check its work, and planning on large changes, both cost tokens and save more than they spend.
This post is about usage and cost. How each mechanism decides what Claude sees, and what survives compaction, is covered in context engineering for Claude Code.
See what you used
/usage(also/costand/stats) opens with a Session block: tokens by model, split into input, output, cache reads and cache writes, and an estimated cost. Anthropic’s cost documentation (opens in a new tab) says the figure is computed locally at list price and is meant for API users; on a subscription it is not what you are billed.- On Pro, Max, Team and Enterprise plans, the same screen shows your plan usage bars and a breakdown of recent usage by skill, subagent, plugin and MCP server. Press
dorwto switch between the last day and the last week. It covers this machine only, not other devices or claude.ai. - The breakdown also flags any behaviour, such as long context or cache misses, that accounts for 10% or more of recent usage. That flag is usually the fastest answer to “why is this session so expensive?”
/contextshows the window as a coloured grid, with suggestions for context-heavy tools, memory files that have grown too large, and warnings as you near capacity./insightswrites a report on how you use Claude Code across recent sessions: what you work on and where sessions go wrong. The analysis itself counts against your usage.
To keep an eye on it without running a command, put the context percentage in the status line (opens in a new tab). Claude Code passes the script a JSON object that includes context_window.used_percentage:
#!/bin/bash input=$(cat) pct=$(echo "$input" | jq -r '.context_window.used_percentage // 0') echo "context $pct%"
For a team, per-person numbers come from outside the terminal: organisation analytics on Team and Enterprise plans, the Console on API billing, or an OpenTelemetry export on any setup.
What actually spends the tokens
Usage climbs in ways that are not obvious from what you typed. The costs page lists the usual reasons a long session draws more than expected:
- The whole conversation goes with every request. Each tool call is another request carrying its results, so a one-line question at the end of a day-long session still pays to re-read the whole history, at the cached rate if the cache is warm.
- Cache misses. The first message after a break longer than the prompt cache lifetime (opens in a new tab) reprocesses the full context. The lifetime is an hour on a subscription and five minutes by default on an API key.
- File reads and command output. A test run that prints a thousand lines, or a search that opens forty files, lands in the window and stays there.
- Thinking. Extended thinking is on by default and thinking tokens are billed as output tokens.
- Subagents, workflows and agent teams. Each sends its own requests on top of the main conversation, and each teammate keeps spending until it exits.
- Scheduled tasks. A
/loopfires on its interval even while you are away, sending the full context each time. - Compaction.
/compacthas to read the conversation to summarise it, so compacting a large context is itself a large request./clearcosts nothing.
Settings that cut usage
- Match the model to the job. Claude Code now starts on Opus on most plans and clouds; drop to Sonnet for routine edits and quick questions to spend less. Switch with
/model, or use theopusplanalias, which runs Opus in plan mode and Sonnet for the edits. Claude Sonnet vs Opus covers which fits which job. - Lower the effort on simple work.
/effort lowormediumcuts thinking on routine tasks; raise it again when the problem is hard. The model configuration guide (opens in a new tab) lists the levels each model supports. - Put subagents on a smaller model. Add
model: haikuto a subagent that only searches or summarises, or setCLAUDE_CODE_SUBAGENT_MODELfor subagents that name no model. The built-in Explore subagent inherits your session’s model; define your ownExplorewithmodel: haikuto keep searches cheap. - Switch off MCP servers you are not using. Tool definitions are deferred by default, so an idle server costs little, and
/mcp disable <server>disconnects it altogether. Where a command-line tool such asghdoes the job, it adds no tool listing at all. - Keep CLAUDE.md under about 200 lines. It loads in every session, so a workflow described there is paid for even when you are doing something else. Move procedures into skills, which load only when used;
/skillscan sort yours by token cost. - Filter noisy output before Claude sees it. A
PreToolUsehook can rewrite a test command to print only failures, turning tens of thousands of tokens of output into a few hundred.
--- name: log-reader description: Reads long logs and test output and reports only the failures that matter. Use for any output over a few hundred lines. tools: Read, Grep, Glob model: haiku --- Read the file or output you are given. Report each distinct failure once, with the file and line, and nothing else. Do not suggest fixes.
A subagent like that keeps the noise in its own window and returns a short summary to yours. More patterns are in Claude Code subagent examples.
Habits that cut usage
- Run
/clearwhen you switch to unrelated work. Use/renamefirst so you can/resumethe old session if you need it. - Compact with a focus before a long new stretch, for example
/compact keep the migration plan and the failing test names, so the summary keeps what you choose. - Ask side questions with
/btw. The answer never enters the conversation, so it is not re-sent on every later message. - Be specific. “Add input validation to the login function in auth.ts” reads one file; “improve this codebase” reads dozens.
- Stop a wrong turn early with
Escrather than letting Claude finish and then asking it to undo. - After a long break on Pro or Max, accept the offer to resume from a summary instead of the full history.
# Compact instructions When compacting, keep the list of modified files, the test commands, and any error messages still unresolved.
What not to cut
- Verification. Letting Claude run the tests costs tokens; a wrong fix you discover tomorrow costs a whole new session.
- Plan mode on large or unfamiliar changes. A plan you correct before any edit is cheaper than an implementation you throw away.
- The CLAUDE.md lines Claude cannot guess. Strip the build command and Claude spends more tokens finding it than the line ever cost.
- Prompt caching. It can be switched off with an environment variable, but caching is what makes re-sending the conversation affordable. Leave it on.
- History you still need. Mid-task, compact rather than clear;
/clearis for when the old conversation is no longer useful.
Stop paying to re-explain the work
A quieter cost is the start of each session: Claude re-reading files, or you re-typing the background, to work out where things stand. Keep that state outside the session. With fenbs connected over MCP, one CLAUDE.md line, “at the start, call fenbs_get_context, then list the In Progress lane”, gives a new session the AI context notes earlier sessions left and the tasks already under way, in two tool calls. Keep those notes short and true; they load every time, just like CLAUDE.md. A task-tracking workflow for Claude Code shows the rest of the rhythm.
Related
The commands above, with the rest of the set: Claude Code commands cheat sheet. Why long sessions get worse as well as dearer: AI context rot. What fenbs keeps for assistants between sessions: AI context, connected through the Claude Code integration.