# How to reduce Claude Code token usage: what uses tokens and what saves them

Most of the tokens Claude Code spends are context it re-reads, not answers it writes. The settings and habits that cut that, straight from Anthropic's docs.

Published 2026-10-07 by Richard Kaminsky and Mitchell Lipyansky, the co-founders of Poly. Canonical: https://usepoly.co/reduce-claude-code-token-usage
Poly is a multiplayer AI coding workspace: a shared room where your team works with one AI agent, together. Free to start: https://usepoly.co/

**To use fewer tokens in Claude Code, keep the context small and pick a cheaper setup before you start. Anthropic's cost guide puts it simply: "Token costs scale with context size: the more context Claude processes, the more tokens you use." The biggest savings: /clear when you switch to unrelated work ("stale context wastes tokens on every subsequent message"), Sonnet 5.5 instead of the default Opus 5.5 for most coding, a lower /effort level, fewer MCP servers, code-intelligence plugins so Claude reads less, hooks and subagents that keep verbose output out of the conversation, and a CLAUDE.md under 200 lines. Choose your model and effort at the start of a session, because changing them mid-session breaks the prompt cache.**

## Key takeaways

- **Clear, don't drag:** `/clear` between tasks costs nothing; a long, mixed session re-sends everything each turn.
- **Right model:** "Sonnet handles most coding tasks well and costs less than Opus." Opus 5.5 is the default.
- **Thinking is output:** "Thinking tokens are billed as output tokens"; `/effort low` or `medium` cuts them.
- **Keep the cache warm:** switching models, effort or MCP servers mid-session forces Claude Code to re-send uncached context.
- **Measure first:** `/usage` shows cost, cache hit rate and what's spending; `/context` shows what's filling the window.

*Ways to cut Claude Code token use, from Anthropic's cost and caching docs, October 7, 2026.*

| Change | How | Saves |
| --- | --- | --- |
| Start fresh between tasks | /clear | Re-sending stale context |
| Use a cheaper model | /model sonnet; model: haiku for subagents | Price per token |
| Lower thinking | /effort low or medium | Output tokens |
| Fewer MCP servers | Disable in /mcp; prefer CLI tools | Tool listings in context |
| Code-intelligence plugins | /plugin install a language's LSP plugin | File reads and searches |
| Filter verbose output | A hook or a subagent | Tens of thousands of tokens per run |
| Short CLAUDE.md | Under 200 lines; move workflows to skills | Context on every request |
| Don't switch mid-session | Pick model and effort at the start | Cache misses |

## What uses the most tokens in Claude Code?

Every request re-sends the whole conversation: system prompt, tools, CLAUDE.md, every message and every file and command output so far. Prompt caching makes repeated context cheap, but it still counts ([Anthropic](https://code.claude.com/docs/en/costs)). The usual culprits:

- **Long sessions mixing unrelated tasks.**
- **Large file reads and verbose test or log output.**
- **Thinking,** which can be "tens of thousands of tokens per request" and is billed as output.
- **Opus left as the default** for routine work.
- **Many MCP servers,** and repeated corrections after a bad start.
- **Agent teams,** which use "approximately 7x more tokens than standard sessions when teammates run in plan mode."

## How do you reduce Claude Code token usage?

From Anthropic's cost guide, in order ([Anthropic](https://code.claude.com/docs/en/costs)):

1. **Manage context.** `/clear` when switching tasks; `/compact Focus on code samples and API usage` to summarize with a focus; add a "Compact instructions" section to CLAUDE.md.
2. **Choose the right model.** Sonnet for most coding, Opus "for complex architectural decisions or multi-step reasoning," and `model: haiku` for simple subagents.
3. **Cut MCP overhead.** "Prefer CLI tools when available": `gh`, `aws`, `gcloud` and `sentry-cli` cost less context than MCP servers. Switch off servers you don't use in `/mcp` ([Claude Code MCP](/claude-code-mcp)).
4. **Install code-intelligence plugins.** "A single 'go to definition' call replaces what might otherwise be a grep followed by reading multiple candidate files"; for example `/plugin install typescript-lsp@claude-plugins-official` ([Claude Code plugins](/claude-code-plugins)).
5. **Trim output with hooks.** A hook that greps test output for errors can reduce "context from tens of thousands of tokens to hundreds" ([Claude Code hooks](/claude-code-hooks)).
6. **Move instructions into skills.** Skills load only when used; "Aim to keep CLAUDE.md under 200 lines" ([Claude Code skills](/claude-code-skills)).
7. **Lower the effort.** `/effort low` or `medium`; Opus, Sonnet and Haiku 5.5 already default to medium.
8. **Delegate verbose work to subagents.** Their output stays in their context and only a summary comes back, though their requests still count.
9. **Write specific prompts.** "Vague requests like 'improve this codebase' trigger broad scanning."
10. **Plan big changes first** and correct early with Esc or `/rewind` ([Claude Code plan mode](/claude-code-plan-mode)).

## How does prompt caching affect Claude Code costs?

Claude Code caches the repeated part of each request automatically; cache reads cost about 5% of normal input, $0.20 per million tokens on Opus 5.5 and $0.10 on Sonnet 5.5 ([Anthropic](https://code.claude.com/docs/en/prompt-caching)). On a Claude subscription, the main conversation's cache lasts an hour; with an API key it's five minutes.

**What breaks the cache:** switching models ("each model has its own cache," and `opusplan` switches on every plan-mode toggle), turning on fast mode, adding or removing an MCP server, compacting, and upgrading Claude Code. Editing CLAUDE.md doesn't break it, but the change only applies after `/clear` or a restart. Anthropic's tip: "Pick your model and effort level at the top of a session, then save /compact for natural breaks between tasks."

## Which model uses the fewest tokens?

They use similar numbers of tokens; they cost different amounts per token. Per million input and output tokens: Opus 5.5 $4 and $20, Sonnet 5.5 $2 and $10, Haiku 5.5 $0.10 and $0.50, rising to $0.50 and $2.50 on prompts over 100,000 tokens, which long Claude Code sessions often pass ([Sonnet 5.5 vs Opus 5.5](/sonnet-5-5-vs-opus-5-5)). Set a cheaper model for subagents with `model: haiku` in their frontmatter, or `CLAUDE_CODE_SUBAGENT_MODEL`. On Pro and Max, tokens count against your five-hour and weekly limits rather than dollars, and those limits are shared across models ([Claude Code usage limits](/claude-code-usage-limits)).

## How do you see how many tokens Claude Code is using?

- **`/usage`:** session cost at list price, the prompt cache hit rate, and on paid plans what's using your allowance, by skill, subagent, plugin and MCP server.
- **`/context`:** what's filling the window, with suggestions ([Claude Code context window](/claude-code-context-window)).
- **Status line:** `cost.total_cost_usd` (an estimate at list price), `context_window.used_percentage` and, on Pro and Max, your five-hour and weekly limits ([Claude Code status line](/claude-code-statusline)).
- **OpenTelemetry:** `claude_code.cost.usage` and `claude_code.token.usage` for a team dashboard.
- **API spend limits:** set a cap on the Console workspace Claude Code creates.

Anthropic's average for teams is about $13 per developer per active day ([Claude Code pricing](/claude-code-pricing)).

## What doesn't work?

- `MAX_THINKING_TOKENS` is ignored on Opus, Sonnet and Haiku 5.5 and the Fable models, which use adaptive reasoning; use `/effort` instead.
- Thinking can't be turned off on those models.
- `@imports` in CLAUDE.md don't save tokens; imported files load at the start anyway.
- Disabling prompt caching makes every request cost more. "For normal use, leave caching enabled."

## And Codex?

The same ideas apply: `/compact` to summarize, `/status` to see remaining context, GPT-6 Luna for focused, repeatable tasks, the default reasoning effort before raising it, a short AGENTS.md and fewer MCP servers ([OpenAI](https://learn.chatgpt.com/docs/pricing)). More in [Codex pricing](/codex-pricing).

## Seeing cost per turn as a team

In Poly (usepoly.co), a team works with one Claude Code or Codex agent in a shared browser room, and every turn's receipt shows who ran it, which model, how full the context is and what it cost, so a team can see where its tokens go. Free to start. [What is Poly?](/what-is-poly)

## Common questions

**How do I reduce token usage in Claude Code?**

Use /clear between unrelated tasks, Sonnet instead of Opus for most coding, a lower /effort level, fewer MCP servers, code-intelligence plugins, hooks or subagents for verbose output, and a CLAUDE.md under 200 lines. Pick your model at the start so you don't break the prompt cache.

**Why does Claude Code use so many tokens?**

Every request re-sends the whole conversation, including CLAUDE.md, tool definitions, files read and command output, and thinking is billed as output. Long sessions that mix tasks and large outputs grow fastest.

**Does /compact save tokens?**

It shrinks the conversation for later requests, but the compaction is itself a large request, and it costs most when you resume an old session. /clear costs nothing and is better when you're switching to unrelated work.

**How can I see my Claude Code token usage?**

Run /usage for session cost, cache hit rate and what's using your allowance, or /context for what's filling the window. A custom status line can show cost and context continuously.

**Does switching models in Claude Code waste tokens?**

It can. Each model has its own prompt cache, so switching mid-session makes the next request re-send uncached context. Anthropic recommends picking the model and effort level at the start of a session.

More guides: https://usepoly.co/guides · Security: https://usepoly.co/security · Pricing: https://usepoly.co/pricing
