
The Claude Code logo. Image: Anthropic
Claude Code burns through your usage because it sends the whole conversation back to Claude with every request, so a long session, a big model, lots of thinking and a crowd of tools all multiply the cost of each message. The fix is to keep the context small, use the cheapest model that can do the job, and know where your usage is going. Anthropic’s own cost guide for Claude Code lays out each habit below, and most of them are one command.
Why is Claude Code using so many tokens?
Because every message carries the full history. Anthropic says Claude Code “sends your full conversation with every request,” and each time Claude uses a tool it sends another request with that batch of results attached. Prompt caching means the old history is re-read at a cheaper rate, but it still counts. In Anthropic’s words, “a one-line question in a session that has been open all day still draws usage for the whole conversation.”
That is why a short chat barely registers and a six-hour session can empty your allowance. Plan usage is also counted against two allowances at once: the rolling five-hour session limit and the weekly limit. Anthropic notes that one burst of heavy activity, such as a large workflow fanning out into many agents, can use up the weekly allowance before the five-hour window has even reset.
For scale, Anthropic says the average across enterprise deployments is about $13 per developer per active day, or $150 to $250 a month, with costs under $30 a day for 90% of users. Those are API-rate figures, but they show how much your habits matter.
How to see what is using your Claude Code usage
Start by measuring. Three built-in tools show where it goes:
/usage: on Pro, Max, Team and Enterprise plans it shows your plan usage bars and a breakdown of recent usage by skills, subagents, plugins and individual MCP servers. It also flags behaviours such as long context or cache misses when one accounts for 10% or more of recent usage, with a tip for each. Pressdorwto switch between the last 24 hours and the last 7 days. The figures are approximate and only cover that machine./context: shows what is filling your context window right now.- A status line: Anthropic’s status line docs list
rate_limits.five_hour.used_percentageandrate_limits.seven_day.used_percentage, so you can keep both limits on screen all the time.
The /usage breakdown also lists your heaviest /loop and scheduled tasks, which is often where a mystery drain turns up.
Quick fixes at a glance
| Habit | How | Why it saves usage |
|---|---|---|
| Start fresh between tasks | /clear | Stale context is re-sent with every later message |
| Use Sonnet, not Opus, for most work | /model | Anthropic says Sonnet handles most coding tasks and costs less |
| Lower the effort level | /effort | Thinking tokens are billed as output tokens |
| Switch off unused MCP servers | /mcp | Each server adds to context |
| Plan before you build | Shift+Tab | Avoids expensive re-work down the wrong path |
| Stop a bad run early | Esc, then /rewind | Stops spending on a wrong direction |
| Keep CLAUDE.md short | Under 200 lines, specifics in skills | CLAUDE.md loads at the start of every session |
All from Anthropic’s Claude Code cost documentation, checked October 7, 2026.
Use /clear when you switch tasks
This is the biggest single saving. Anthropic’s advice is to use /clear when you move to unrelated work, because “stale context wastes tokens on every subsequent message.” If you want to come back later, run /rename first and /resume afterwards.
When you do want continuity, /compact summarises the conversation, and you can steer it, for example /compact Focus on code samples and API usage. Be aware that compacting reads the whole conversation, so Anthropic warns that “compacting a large context is itself a large request”, whereas /clear costs nothing. Compact earlier rather than later, and clear whenever the old context is no use.
Which Claude model uses the least usage in Claude Code?
Anthropic recommends Sonnet for most coding tasks and keeping Opus for “complex architectural decisions or multi-step reasoning.” Switch with /model or set a default in /config. Two catches:
- Switching to Opus also switches your subagents if they inherit your session’s model. For simple subagent jobs, set
model: haikuin the subagent’s configuration. - Each model has its own cache. Anthropic says that after a switch “the next request re-reads the whole conversation with no cache hits”, so pick your model at the start of a task, or after a
/clear, rather than flipping back and forth mid-session.
For the models themselves, see our guides to Claude Sonnet 5.5 and Claude Haiku 5.5, the cheapest current model.
Turn down extended thinking and effort
Extended thinking is on by default, and Anthropic says “thinking tokens are billed as output tokens, and the default budget can be tens of thousands of tokens per request.” For simple jobs, lower the effort with /effort or in /model. The levels are low, medium, high, xhigh and max.
You can’t turn thinking off on Opus 5.5, Sonnet 5.5, Haiku 5.5 or the Fable models, which always use it, so effort is your only dial there. The MAX_THINKING_TOKENS setting only works on models with a fixed thinking budget, and adaptive models ignore it.
Trim MCP servers, big files and CLAUDE.md
- MCP servers: tool definitions are deferred by default, so only names enter context until a tool is used, but unused servers still add overhead. Run
/mcpto disable any you aren’t using. Anthropic also says command-line tools such asgh,awsandgcloudare more context-efficient than MCP servers. - CLAUDE.md: it loads at the start of every session, so detailed workflow instructions sit in context even when you’re doing something unrelated. Anthropic suggests moving them into skills, which load only when needed, and keeping CLAUDE.md under 200 lines.
- Noisy output: a hook can filter a 10,000-line log down to the error lines before Claude reads it, cutting “tens of thousands of tokens to hundreds.” Anthropic’s page has a ready-made test-output filter.
- Typed languages: code intelligence plugins let Claude jump to a definition instead of searching and reading several files.
Watch subagents, agent teams and loops
Helpers draw on the same pot. Anthropic says each subagent “sends its own requests on top of the main conversation’s”, and agent teams use roughly 7x more tokens than a standard session when teammates run in plan mode, because each one keeps its own context window. Teammates keep consuming until they exit.
Subagents still help when used for noisy jobs such as running tests or reading logs, since the long output stays in their context and only a summary comes back. Pick a smaller model for them, and keep team tasks small and self-contained. The /usage attribution list shows how much of your usage the subagents took.
Why Claude Code uses usage when you’re not typing
A session left open can keep spending. Anthropic lists the causes:
- Cache misses after a break. The cache lasts an hour on a subscription, so your first message after a longer pause reprocesses the full context. On Pro and Max, Claude Code offers to resume a big session from a summary instead.
- Scheduled tasks and loops fire on their interval even when the session is idle, sending your full context each time.
- Messages from other sessions arrive as new turns. Set
crossSessionInboundtoholdto stop that. - Goal check-ins start up to three idle turns per goal. Set
CLAUDE_CODE_GOAL_CHECKIN_MINUTESto0to switch them off. - Background jobs such as summaries for
claude --resumeuse little, typically under $0.04 per session. Prompt suggestions also send a short request after each reply, which you can turn off.
Running several sessions at once, including cloud sessions, uses your limits faster in proportion, since they all draw from the same allowance.
Prompt smarter to avoid wasted runs
- Be specific. Anthropic contrasts “improve this codebase”, which triggers broad scanning, with “add input validation to the login function in auth.ts”.
- Use plan mode (Shift+Tab) for complex work so Claude proposes an approach before it builds.
- Course-correct early. Press Esc the moment it goes the wrong way, or double-tap it or use
/rewindto restore an earlier checkpoint. - Give it a way to check itself, such as test cases or the expected output, and test one file at a time.
How to see your limit coming
Claude Code can warn you before the window closes, with a message such as “You’ve used 85% of your session limit”, and the message shows when it resets. With the status line fields above you can watch both the five-hour and weekly bars all day. If you do hit a limit, the session and weekly allowances are shared across all models, so switching models won’t restore access, but the separate Opus and Sonnet limits can be dodged by moving to the other family. Our guides to Claude Code’s wrap-up allowance and the one-time “Reset for free” (it expires October 22, 2026) cover what to do next.
Limits work the same worldwide, and prices are in US dollars before local tax.
Last checked: October 7, 2026, against Anthropic’s Claude Code documentation.
Frequently asked questions
Why is Claude Code using so many tokens?
Claude Code sends your full conversation with every request, and again each time Claude uses a tool, so long sessions, big files, lots of thinking, MCP servers and subagents all add up. Anthropic says even a one-line question in a session open all day draws usage for the whole conversation.
How do I check how much Claude Code usage I have used?
Run /usage. On Pro, Max, Team and Enterprise plans it shows your plan usage bars and a breakdown by skills, subagents, plugins and MCP servers, with the last 24 hours or 7 days switchable. Run /context to see what is filling the context window.
Does /clear save tokens in Claude Code?
Yes. Anthropic says stale context wastes tokens on every later message, and /clear starts fresh at no cost. Run /rename first if you want to /resume the old session later.
Does /compact save usage?
It shrinks the context for later messages, but compacting reads the whole conversation, so compacting a large context is itself a large request. Compact earlier, or use /clear when you do not need the old context.
Is Sonnet cheaper than Opus in Claude Code?
Yes. Anthropic says Sonnet handles most coding tasks and costs less than Opus, which it suggests keeping for complex architecture and multi-step reasoning. Switch with /model.
Does extended thinking use more tokens?
Yes. Thinking tokens are billed as output tokens, and the default budget can be tens of thousands per request. Lower the effort with /effort. Thinking cannot be switched off on Opus 5.5, Sonnet 5.5, Haiku 5.5 or the Fable models.
Do MCP servers use up Claude Code tokens?
They add overhead. Tool definitions are deferred by default, so only names enter context until a tool is used, but Anthropic still recommends disabling servers you are not using with /mcp and preferring command-line tools such as gh, aws and gcloud.
Does Claude Code use usage when idle?
It can. Scheduled tasks, loops, messages from other sessions and goal check-ins can start turns in an idle session and send your full context each time, and the first message after a break longer than the cache lifetime (an hour on a subscription) reprocesses everything.
How often do Claude Code limits reset?
The session limit resets on a rolling five-hour window and there is a separate weekly limit. The message you see when you hit one shows the reset time, and /usage shows when each resets.
Sources: Anthropic’s Claude Code documentation (manage costs effectively, error reference, status line, model configuration).





