Guide
How to reduce Claude Code token usage: what uses tokens and what saves them
Most of the tokens Claude Code spends are context it re-reads, not answers it writes. The settings and habits that cut that, straight from Anthropic's docs.
To use fewer tokens in Claude Code, keep the context small and pick a cheaper setup before you start. Anthropic's cost guide puts it simply: "Token costs scale with context size: the more context Claude processes, the more tokens you use." The biggest savings: /clear when you switch to unrelated work ("stale context wastes tokens on every subsequent message"), Sonnet 5.5 instead of the default Opus 5.5 for most coding, a lower /effort level, fewer MCP servers, code-intelligence plugins so Claude reads less, hooks and subagents that keep verbose output out of the conversation, and a CLAUDE.md under 200 lines. Choose your model and effort at the start of a session, because changing them mid-session breaks the prompt cache.
Key takeaways
- Clear, don't drag:
/clearbetween tasks costs nothing; a long, mixed session re-sends everything each turn. - Right model: "Sonnet handles most coding tasks well and costs less than Opus." Opus 5.5 is the default.
- Thinking is output: "Thinking tokens are billed as output tokens";
/effort lowormediumcuts them. - Keep the cache warm: switching models, effort or MCP servers mid-session forces Claude Code to re-send uncached context.
- Measure first:
/usageshows cost, cache hit rate and what's spending;/contextshows what's filling the window.
| Change | How | Saves |
|---|---|---|
| Start fresh between tasks | /clear | Re-sending stale context |
| Use a cheaper model | /model sonnet; model: haiku for subagents | Price per token |
| Lower thinking | /effort low or medium | Output tokens |
| Fewer MCP servers | Disable in /mcp; prefer CLI tools | Tool listings in context |
| Code-intelligence plugins | /plugin install a language's LSP plugin | File reads and searches |
| Filter verbose output | A hook or a subagent | Tens of thousands of tokens per run |
| Short CLAUDE.md | Under 200 lines; move workflows to skills | Context on every request |
| Don't switch mid-session | Pick model and effort at the start | Cache misses |
What uses the most tokens in Claude Code?
Every request re-sends the whole conversation: system prompt, tools, CLAUDE.md, every message and every file and command output so far. Prompt caching makes repeated context cheap, but it still counts (Anthropic). The usual culprits:
- Long sessions mixing unrelated tasks.
- Large file reads and verbose test or log output.
- Thinking, which can be "tens of thousands of tokens per request" and is billed as output.
- Opus left as the default for routine work.
- Many MCP servers, and repeated corrections after a bad start.
- Agent teams, which use "approximately 7x more tokens than standard sessions when teammates run in plan mode."
How do you reduce Claude Code token usage?
From Anthropic's cost guide, in order (Anthropic):
- Manage context.
/clearwhen switching tasks;/compact Focus on code samples and API usageto summarize with a focus; add a "Compact instructions" section to CLAUDE.md. - Choose the right model. Sonnet for most coding, Opus "for complex architectural decisions or multi-step reasoning," and
model: haikufor simple subagents. - Cut MCP overhead. "Prefer CLI tools when available":
gh,aws,gcloudandsentry-clicost less context than MCP servers. Switch off servers you don't use in/mcp(Claude Code MCP). - Install code-intelligence plugins. "A single 'go to definition' call replaces what might otherwise be a grep followed by reading multiple candidate files"; for example
/plugin install typescript-lsp@claude-plugins-official(Claude Code plugins). - Trim output with hooks. A hook that greps test output for errors can reduce "context from tens of thousands of tokens to hundreds" (Claude Code hooks).
- Move instructions into skills. Skills load only when used; "Aim to keep CLAUDE.md under 200 lines" (Claude Code skills).
- Lower the effort.
/effort lowormedium; Opus, Sonnet and Haiku 5.5 already default to medium. - Delegate verbose work to subagents. Their output stays in their context and only a summary comes back, though their requests still count.
- Write specific prompts. "Vague requests like 'improve this codebase' trigger broad scanning."
- Plan big changes first and correct early with Esc or
/rewind(Claude Code plan mode).
How does prompt caching affect Claude Code costs?
Claude Code caches the repeated part of each request automatically; cache reads cost about 5% of normal input, $0.20 per million tokens on Opus 5.5 and $0.10 on Sonnet 5.5 (Anthropic). On a Claude subscription, the main conversation's cache lasts an hour; with an API key it's five minutes.
What breaks the cache: switching models ("each model has its own cache," and opusplan switches on every plan-mode toggle), turning on fast mode, adding or removing an MCP server, compacting, and upgrading Claude Code. Editing CLAUDE.md doesn't break it, but the change only applies after /clear or a restart. Anthropic's tip: "Pick your model and effort level at the top of a session, then save /compact for natural breaks between tasks."
Which model uses the fewest tokens?
They use similar numbers of tokens; they cost different amounts per token. Per million input and output tokens: Opus 5.5 $4 and $20, Sonnet 5.5 $2 and $10, Haiku 5.5 $0.10 and $0.50, rising to $0.50 and $2.50 on prompts over 100,000 tokens, which long Claude Code sessions often pass (Sonnet 5.5 vs Opus 5.5). Set a cheaper model for subagents with model: haiku in their frontmatter, or CLAUDE_CODE_SUBAGENT_MODEL. On Pro and Max, tokens count against your five-hour and weekly limits rather than dollars, and those limits are shared across models (Claude Code usage limits).
How do you see how many tokens Claude Code is using?
/usage: session cost at list price, the prompt cache hit rate, and on paid plans what's using your allowance, by skill, subagent, plugin and MCP server./context: what's filling the window, with suggestions (Claude Code context window).- Status line:
cost.total_cost_usd(an estimate at list price),context_window.used_percentageand, on Pro and Max, your five-hour and weekly limits (Claude Code status line). - OpenTelemetry:
claude_code.cost.usageandclaude_code.token.usagefor a team dashboard. - API spend limits: set a cap on the Console workspace Claude Code creates.
Anthropic's average for teams is about $13 per developer per active day (Claude Code pricing).
What doesn't work?
MAX_THINKING_TOKENSis ignored on Opus, Sonnet and Haiku 5.5 and the Fable models, which use adaptive reasoning; use/effortinstead.- Thinking can't be turned off on those models.
@importsin CLAUDE.md don't save tokens; imported files load at the start anyway.- Disabling prompt caching makes every request cost more. "For normal use, leave caching enabled."
And Codex?
The same ideas apply: /compact to summarize, /status to see remaining context, GPT-6 Luna for focused, repeatable tasks, the default reasoning effort before raising it, a short AGENTS.md and fewer MCP servers (OpenAI). More in Codex pricing.
Seeing cost per turn as a team
In Poly (usepoly.co), a team works with one Claude Code or Codex agent in a shared browser room, and every turn's receipt shows who ran it, which model, how full the context is and what it cost, so a team can see where its tokens go. Free to start. What is Poly?
Common questions
How do I reduce token usage in Claude Code?
Use /clear between unrelated tasks, Sonnet instead of Opus for most coding, a lower /effort level, fewer MCP servers, code-intelligence plugins, hooks or subagents for verbose output, and a CLAUDE.md under 200 lines. Pick your model at the start so you don't break the prompt cache.
Why does Claude Code use so many tokens?
Every request re-sends the whole conversation, including CLAUDE.md, tool definitions, files read and command output, and thinking is billed as output. Long sessions that mix tasks and large outputs grow fastest.
Does /compact save tokens?
It shrinks the conversation for later requests, but the compaction is itself a large request, and it costs most when you resume an old session. /clear costs nothing and is better when you're switching to unrelated work.
How can I see my Claude Code token usage?
Run /usage for session cost, cache hit rate and what's using your allowance, or /context for what's filling the window. A custom status line can show cost and context continuously.
Does switching models in Claude Code waste tokens?
It can. Each model has its own prompt cache, so switching mid-session makes the next request re-send uncached context. Anthropic recommends picking the model and effort level at the start of a session.