# Claude Code context window: how big it is, what fills it, and what to do when it's full

A million tokens sounds endless and isn't. What takes up the room, how to see it, and the commands that clear it without losing your place.

Published 2026-10-07 by Richard Kaminsky and Mitchell Lipyansky, the co-founders of Poly. Canonical: https://usepoly.co/claude-code-context-window
Poly is a multiplayer AI coding workspace: a shared room where your team works with one AI agent, together. Free to start: https://usepoly.co/

**Claude Code's current models, Opus 5.5 (the default), Sonnet 5.5, Fable 5.1 and Haiku 5.5, have a context window of 1 million tokens on every plan, Pro included, with up to 128,000 tokens of output per response; there's no extra charge for using the long window. The window holds everything Claude is working with: its instructions, your CLAUDE.md, tool definitions, the conversation, and every file it has read. Type /context to see what's filling it. When it nears the limit, Claude Code compacts the conversation automatically, at about 967,000 tokens by default; you can run /compact yourself, or /clear to start fresh. A full context window is not the same as hitting your plan's usage limit.**

## Key takeaways

- **Size:** 1 million tokens on Opus 5.5, Sonnet 5.5, Fable 5.1 and Haiku 5.5, on every plan; the older Haiku 4.5 has 200,000.
- **See it:** `/context` shows a grid of what's using the window; `/usage` shows cost and plan limits.
- **Auto-compact** at about 967,000 tokens; change it with `/autocompact`, or turn it off in `/config`.
- **`/compact`** summarizes and keeps going; **`/clear`** starts over and costs nothing.
- **Why it matters:** Anthropic says "performance degrades as it fills," so a fresh context often beats a long one.

*Context windows in Claude Code, from Anthropic's docs, October 7, 2026.*

| Model | Context window | Max output | Long-context price |
| --- | --- | --- | --- |
| Opus 5.5 (default) | 1M | 128K | Same rate at any length |
| Sonnet 5.5 | 1M | 128K | Same rate at any length |
| Fable 5.1 | 1M | 128K | Same rate at any length |
| Haiku 5.5 (new) | 1M | 128K | Higher above 100K tokens |
| Haiku 4.5 | 200K | 64K | Not applicable |

## How big is Claude Code's context window?

From Anthropic's model docs ([Anthropic](https://code.claude.com/docs/en/model-config)):

- **Opus 5.5, Sonnet 5.5 and Fable 5.1:** 1 million tokens, native. Models from Opus 4.7 on "run with the 1M window on every plan, including Pro." There's no `[1m]` variant to choose and no usage credits needed.
- **Haiku 5.5,** released October 7, 2026 and now the default Haiku: 1 million tokens.
- **Haiku 4.5:** 200,000 tokens and 64,000 of output.
- **Output:** up to 128,000 tokens per response on the 1M models.

On the API there's no long-context surcharge: "a 900k-token request is billed at the same per-token rate as a 9k-token request," except Haiku 5.5, which costs more above 100,000 tokens ([Anthropic](https://platform.claude.com/docs/en/about-claude/pricing)). To keep sessions at 200,000 tokens, set `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`.

## What fills the context window?

Everything sent with each request counts ([Anthropic](https://code.claude.com/docs/en/context-window)):

- **At the start:** Claude Code's system prompt and tools, your CLAUDE.md files, the first part of auto memory, skill descriptions (capped at 1% of the window) and the names of MCP tools.
- **As you work:** your messages, Claude's replies and thinking, every file it reads and every command's output.
- **On demand:** MCP tool schemas load only when used, because tool search is on by default ([Claude Code MCP](/claude-code-mcp)); skill bodies load when invoked ([Claude Code skills](/claude-code-skills)).

Large file reads and long command output fill it fastest. Hooks cost nothing unless they return output ([Claude Code hooks](/claude-code-hooks)).

## How do you check context usage in Claude Code?

- **`/context`** shows a colored grid of what's using the window (system prompt, tools, MCP tools, memory files, skills and messages) with suggestions when something is heavy ([Anthropic](https://code.claude.com/docs/en/commands)).
- **`/usage`** shows session cost and how much of your plan's limits you've used; `/cost` is an alias.
- **The status line** can show it: a custom status line receives `context_window.used_percentage` and the window's size ([Anthropic](https://code.claude.com/docs/en/statusline), [Claude Code status line](/claude-code-statusline)).

## What happens when the context window is full?

Claude Code compacts automatically: it summarizes the conversation so far and carries on. On the 1M models that happens at about 967,000 tokens by default ([Anthropic](https://code.claude.com/docs/en/model-config)).

- **Compact earlier:** `/autocompact 500k` (anything from 100k to 1M), or the `autoCompactWindow` setting.
- **Turn it off:** switch Auto-compact off in `/config`, or set `DISABLE_AUTO_COMPACT=1`. When you hit the limit you'll see "Context limit reached · /compact or /clear to continue."

## /compact vs /clear vs rewind

- **`/compact`** summarizes the conversation and keeps going. Steer it: `/compact focus on the auth bug fix`. It's a large request itself.
- **`/clear`** (also `/new` or `/reset`) starts a new conversation with an empty context. Anthropic notes "/clear costs nothing."
- **Rewind** (Esc twice, or `/rewind`) goes back to an earlier point and can summarize from there.

**What survives compaction:** the system prompt, your project's CLAUDE.md and auto memory (re-read from disk), up to five recently read files, and skills you invoked, "capped at 5,000 tokens per skill and 25,000 tokens total." Details from early in a long session can be lost, which is why Anthropic suggests putting compaction instructions in CLAUDE.md.

## How do you manage context in Claude Code?

Anthropic's best practices start from this: "Most best practices are based on one constraint: Claude's context window fills up fast, and performance degrades as it fills" ([Anthropic](https://code.claude.com/docs/en/best-practices)).

- **`/clear` between unrelated tasks.**
- **Use subagents for investigation:** they read the files in their own context and return a summary ([Claude Code subagents](/claude-code-subagents)).
- **Be specific** about which files and functions matter, so Claude reads less.
- **Start over after two failed corrections:** "/clear and write a better initial prompt."
- **Keep CLAUDE.md short:** it's read on every request.

More in [Claude Code best practices](/claude-code-best-practices) and [how to reduce Claude Code token usage](/reduce-claude-code-token-usage).

## Is a full context window the same as a usage limit?

No. Anthropic's cost guide lists "a context or auto-compact warning" as "not a usage limit." Usage limits are your plan's five-hour and weekly allowances, shared across models and with Claude chat; a long conversation does use more of them, because each request re-sends the whole context ([Claude Code usage limits](/claude-code-usage-limits)).

## What about Codex?

OpenAI's GPT-6.1 Sol has a 1.05 million-token window on the API, with input above 272,000 tokens charged at twice the rate ([OpenAI](https://developers.openai.com/api/docs/models/gpt-6.1-sol)). In Codex, `/status` shows remaining context and `/compact` summarizes ([OpenAI](https://learn.chatgpt.com/docs/developer-commands)); OpenAI's docs don't state the window Codex uses. More in [How to use Codex](/how-to-use-codex).

## Context in a shared room

In Poly (usepoly.co), a team works with one Claude Code or Codex agent in a shared browser room, and every turn's receipt shows how full the context is, so everyone can see when it's time to start fresh. Free to start. [What is Poly?](/what-is-poly)

## Common questions

**How big is Claude Code's context window?**

1 million tokens on Opus 5.5, Sonnet 5.5, Fable 5.1 and Haiku 5.5, on every plan including Pro, with up to 128,000 tokens of output per response. The older Haiku 4.5 has 200,000.

**How do I see how much context Claude Code is using?**

Type /context for a grid of what's filling the window. /usage shows cost and plan limits, and a custom status line can show the percentage used.

**What happens when Claude Code's context is full?**

It compacts automatically, summarizing the conversation, at about 967,000 tokens by default on the 1M models. You can run /compact yourself, change the threshold with /autocompact, or /clear to start fresh.

**What is the difference between /compact and /clear?**

/compact summarizes the conversation and keeps going, which is itself a large request. /clear starts a new conversation with empty context and costs nothing.

**Does a full context window mean I hit my usage limit?**

No. A context or auto-compact warning is not a usage limit. Usage limits are your plan's five-hour and weekly allowances, though long conversations use them faster because each request re-sends the context.

More guides: https://usepoly.co/guides · Security: https://usepoly.co/security · Pricing: https://usepoly.co/pricing
