← Back to daily report
+78 lines added
-44 lines removed
# Manage costs effectively¶
¶
> Learn how to track and optimize token usage and costs when using Claude Code.¶
¶
Claude Code consumes tokens for each interaction. The average cost is \$6 per developer per day, with daily costs remaining below \$12 for 90% of users.¶
¶
For team usage, Claude Code charges by API token consumption. On average, Claude Code costs \~\$100-200/developer per month with Sonnet 4.5 though there is large variance depending on how many instances users are running and whether they're using it in automation.¶
¶
## Track your costs¶
¶
### Using the `/cost` command¶
¶
<Note>¶
The `/cost` command is not intended for Claude Max and Pro subscribers.¶
</Note>¶
¶
The `/cost` command provides detailed token usage statistics for your current session:¶
¶
```¶
Total cost: $0.55¶
Total duration (API): 6m 19.7s¶
Total duration (wall): 6h 33m 10.2s¶
Total code changes: 0 lines added, 0 lines removed¶
```¶
¶
### Additional tracking options¶
¶
Check [historical usage](https://support.claude.com/en/articles/9534590-cost-and-usage-reporting-in-console) in the Claude Console (requires Admin or Billing role) and set [workspace spend limits](https://support.claude.com/en/articles/9796807-creating-and-managing-workspaces) for the Claude Code workspace (requires Admin role).¶
¶
<Note>¶
When you first authenticate Claude Code with your Claude Console account, a workspace called "Claude Code" is automatically created for you. This workspace provides centralized cost tracking and management for all Claude Code usage in your organization. You cannot create API keys for this workspace - it is exclusively for Claude Code authentication and usage.¶
</Note>¶
¶
## Managing costs for teams¶
¶
When using Claude API, you can limit the total Claude Code workspace spend. To configure, [follow these instructions](https://support.claude.com/en/articles/9796807-creating-and-managing-workspaces). Admins can view cost and usage reporting by [following these instructions](https://support.claude.com/en/articles/9534590-cost-and-usage-reporting-in-console).¶
¶
On Bedrock and Vertex, Claude Code does not send metrics from your cloud. In order to get cost metrics, several large enterprises reported using [LiteLLM](/en/third-party-integrations#litellm), which is an open-source tool that helps companies [track spend by key](https://docs.litellm.ai/docs/proxy/virtual_keys#tracking-spend). This project is unaffiliated with Anthropic and we have not audited its security.¶
¶
### Rate limit recommendations¶
¶
When setting up Claude Code for teams, consider these Token Per Minute (TPM) and Request Per Minute (RPM) per-user recommendations based on your organization size:¶
¶
| Team size | TPM per user | RPM per user |¶
| ------------- | ------------ | ------------ |¶
| 1-5 users | 200k-300k | 5-7 |¶
| 5-20 users | 100k-150k | 2.5-3.5 |¶
| 20-50 users | 50k-75k | 1.25-1.75 |¶
| 50-100 users | 25k-35k | 0.62-0.87 |¶
| 100-500 users | 15k-20k | 0.37-0.47 |¶
| 500+ users | 10k-15k | 0.25-0.35 |¶
¶
For example, if you have 200 users, you might request 20k TPM for each user, or 4 million total TPM (200\*20,000 = 4 million).¶
¶
The TPM per user decreases as team size grows because we expect fewer users to use Claude Code concurrently in larger organizations. These rate limits apply at the organization level, not per individual user, which means individual users can temporarily consume more than their calculated share when others aren't actively using the service.¶
¶
<Note>¶
If you anticipate scenarios with unusually high concurrent usage (such as live training sessions with large groups), you may need higher TPM allocations per user.¶
</Note>¶
¶
## Reduce token usage¶
¶
* **Compact conversations:**¶
¶
* Claude uses auto-compact by default when context reaches approximately 95% capacity. To trigger compaction earlier, set [`CLAUDE_AUTOCOMPACT_PCT_OVERRIDE`](/en/settings#environment-variables) to a lower percentage (for example, `50`)¶
* Toggle auto-compact: Run `/config` and navigate to "Auto-compact enabled"¶
* Use `/compact` manually when context gets large¶
* Add custom instructions: `/compact Focus on code samples and API usage`¶
* Customize compaction by adding to CLAUDE.md:¶
¶
```markdown theme={null}¶
# Summary instructions¶
¶
When you are using compact, please focus on test output and code changes¶
```¶
¶
* **Write specific queries:** Avoid vague requests that trigger unnecessary scanning¶
¶
* **Break down complex tasks:** Split large tasks into focused interactions¶
¶
* **Clear history between tasks:** Use `/clear` to reset context¶
¶
Costs can vary significantly based on:¶
¶
* Size of codebase being analyzed¶
* Complexity of queries¶
* Number of files being searched or modified¶
* Length of conversation history¶
* Frequency of compacting conversations¶
¶
## Background token usage¶
¶
Claude Code uses tokens for some background functionality even when idle:¶
¶
* **Conversation summarization**: Background jobs that summarize previous conversations for the `claude --resume` feature¶
* **Command processing**: Some commands like `/cost` may generate requests to check status¶
¶
These background processes consume a small amount of tokens (typically under \$0.04 per session) even without active interaction.¶
¶
## Tracking version changes and updates¶
¶
### Current version information¶
¶
To check your current Claude Code version and installation details:¶
¶
```bash theme={null}¶
claude doctor¶
```¶
¶
This command shows your version, installation type, and system information.¶
¶
### Understanding changes in Claude Code behavior¶
¶
Claude Code regularly receives updates that may change how features work, including cost reporting:¶
¶
* **Version tracking**: Use `claude doctor` to see your current version¶
* **Behavior changes**: Features like `/cost` may display information differently across versions¶
* **Documentation access**: Claude always has access to the latest documentation, which can help explain current feature behavior¶
¶
### When cost reporting changes¶
¶
If you notice changes in how costs are displayed (such as the `/cost` command showing different information):¶
¶
1. **Verify your version**: Run `claude doctor` to confirm your current version¶
2. **Consult documentation**: Ask Claude directly about current feature behavior, as it has access to up-to-date documentation¶
3. **Contact support**: For specific billing questions, contact Anthropic support through your Console account¶
¶
<Note>¶
For team deployments, we recommend starting with a small pilot group to¶
establish usage patterns before wider rollout.¶
</Note>Track token usage, set team spend limits, and reduce Claude Code costs with context management, model selection, extended thinking settings, and preprocessing hooks.¶
¶
Claude Code consumes tokens for each interaction. Costs vary based on codebase size, query complexity, and conversation length. The average cost is \$6 per developer per day, with daily costs remaining below \$12 for 90% of users.¶
¶
For team usage, Claude Code charges by API token consumption. On average, Claude Code costs \~\$100-200/developer per month with Sonnet 4.5 though there is large variance depending on how many instances users are running and whether they're using it in automation.¶
¶
This page covers how to [track your costs](#track-your-costs), [manage costs for teams](#managing-costs-for-teams), and [reduce token usage](#reduce-token-usage).¶
¶
## Track your costs¶
¶
### Using the `/cost` command¶
¶
<Note>¶
The `/cost` command shows API token usage and is intended for API users. Claude Max and Pro subscribers have usage included in their subscription, so `/cost` data isn't relevant for billing purposes. Subscribers can use `/stats` to view usage patterns.¶
</Note>¶
¶
The `/cost` command provides detailed token usage statistics for your current session:¶
¶
```¶
Total cost: $0.55¶
Total duration (API): 6m 19.7s¶
Total duration (wall): 6h 33m 10.2s¶
Total code changes: 0 lines added, 0 lines removed¶
```¶
¶
## Managing costs for teams¶
¶
When using Claude API, you can [set workspace spend limits](https://platform.claude.com/docs/en/build-with-claude/workspaces#workspace-limits) on the total Claude Code workspace spend. Admins can [view cost and usage reporting](https://platform.claude.com/docs/en/build-with-claude/workspaces#usage-and-cost-tracking) in the Console.¶
¶
<Note>¶
When you first authenticate Claude Code with your Claude Console account, a workspace called "Claude Code" is automatically created for you. This workspace provides centralized cost tracking and management for all Claude Code usage in your organization. You cannot create API keys for this workspace; it is exclusively for Claude Code authentication and usage.¶
</Note>¶
¶
On Bedrock, Vertex, and Foundry, Claude Code does not send metrics from your cloud. To get cost metrics, several large enterprises reported using [LiteLLM](/en/llm-gateway#litellm-configuration), which is an open-source tool that helps companies [track spend by key](https://docs.litellm.ai/docs/proxy/virtual_keys#tracking-spend). This project is unaffiliated with Anthropic and we have not audited its security.¶
¶
### Rate limit recommendations¶
¶
When setting up Claude Code for teams, consider these Token Per Minute (TPM) and Request Per Minute (RPM) per-user recommendations based on your organization size:¶
¶
| Team size | TPM per user | RPM per user |¶
| ------------- | ------------ | ------------ |¶
| 1-5 users | 200k-300k | 5-7 |¶
| 5-20 users | 100k-150k | 2.5-3.5 |¶
| 20-50 users | 50k-75k | 1.25-1.75 |¶
| 50-100 users | 25k-35k | 0.62-0.87 |¶
| 100-500 users | 15k-20k | 0.37-0.47 |¶
| 500+ users | 10k-15k | 0.25-0.35 |¶
¶
For example, if you have 200 users, you might request 20k TPM for each user, or 4 million total TPM (200\*20,000 = 4 million).¶
¶
The TPM per user decreases as team size grows because we expect fewer users to use Claude Code concurrently in larger organizations. These rate limits apply at the organization level, not per individual user, which means individual users can temporarily consume more than their calculated share when others aren't actively using the service.¶
¶
<Note>¶
If you anticipate scenarios with unusually high concurrent usage (such as live training sessions with large groups), you may need higher TPM allocations per user.¶
</Note>¶
¶
## Reduce token usage¶
¶
Token costs scale with context size: the more context Claude processes, the more tokens you use. Claude Code automatically optimizes costs through prompt caching (which reduces costs for repeated content like system prompts) and auto-compaction (which summarizes conversation history when approaching context limits).¶
¶
The following strategies help you keep context small and reduce per-message costs.¶
¶
### Manage context proactively¶
¶
Use `/cost` to check your current token usage, or [configure your status line](/en/statusline#context-window-usage) to display it continuously.¶
¶
* **Clear between tasks**: Use `/clear` to start fresh when switching to unrelated work. Stale context wastes tokens on every subsequent message. Use `/rename` before clearing so you can easily find the session later, then `/resume` to return to it.¶
* **Add custom compaction instructions**: `/compact Focus on code samples and API usage` tells Claude what to preserve during summarization.¶
¶
You can also customize compaction behavior in your CLAUDE.md:¶
¶
```markdown theme={null}¶
# Compact instructions¶
¶
When you are using compact, please focus on test output and code changes¶
```¶
¶
### Choose the right model¶
¶
Sonnet handles most coding tasks well and costs less than Opus. Reserve Opus for complex architectural decisions or multi-step reasoning. Use `/model` to switch models mid-session, or set a default in `/config`. For simple subagent tasks, specify `model: haiku` in your [subagent configuration](/en/sub-agents#choose-a-model).¶
¶
### Reduce MCP server overhead¶
¶
Each MCP server adds tool definitions to your context, even when idle. Run `/context` to see what's consuming space.¶
¶
* **Prefer CLI tools when available**: Tools like `gh`, `aws`, `gcloud`, and `sentry-cli` are more context-efficient than MCP servers because they don't add persistent tool definitions. Claude can run CLI commands directly without the overhead.¶
* **Disable unused servers**: Run `/mcp` to see configured servers and disable any you're not actively using.¶
* **Tool search is automatic**: When MCP tool descriptions exceed 10% of your context window, Claude Code automatically defers them and loads tools on-demand via [tool search](/en/mcp#scale-with-mcp-tool-search). Since deferred tools only enter context when actually used, a lower threshold means fewer idle tool definitions consuming space. Set a lower threshold with `ENABLE_TOOL_SEARCH=auto:<N>` (for example, `auto:5` triggers when tools exceed 5% of your context window).¶
¶
### Offload processing to hooks and skills¶
¶
Custom [hooks](/en/hooks) can preprocess data before Claude sees it. Instead of Claude reading a 10,000-line log file to find errors, a hook can grep for `ERROR` and return only matching lines, reducing context from tens of thousands of tokens to hundreds.¶
¶
A [skill](/en/skills) can give Claude domain knowledge so it doesn't have to explore. For example, a "codebase-overview" skill could describe your project's architecture, key directories, and naming conventions. When Claude invokes the skill, it gets this context immediately instead of spending tokens reading multiple files to understand the structure.¶
¶
For example, this PreToolUse hook filters test output to show only failures:¶
¶
<Tabs>¶
<Tab title="settings.json">¶
Add this to your [settings.json](/en/settings#settings-files) to run the hook before every Bash command:¶
¶
```json theme={null}¶
{¶
"hooks": {¶
"PreToolUse": [¶
{¶
"matcher": "Bash",¶
"hooks": [¶
{¶
"type": "command",¶
"command": "~/.claude/hooks/filter-test-output.sh"¶
}¶
]¶
}¶
]¶
}¶
}¶
```¶
</Tab>¶
¶
<Tab title="filter-test-output.sh">¶
The hook calls this script, which checks if the command is a test runner and modifies it to show only failures:¶
¶
```bash theme={null}¶
#!/bin/bash¶
input=$(cat)¶
cmd=$(echo "$input" | jq -r '.tool_input.command')¶
¶
# If running tests, filter to show only failures¶
if [[ "$cmd" =~ ^(npm test|pytest|go test) ]]; then¶
filtered_cmd="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"¶
echo "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"allow\",\"updatedInput\":{\"command\":\"$filtered_cmd\"}}}"¶
else¶
echo "{}"¶
fi¶
```¶
</Tab>¶
</Tabs>¶
¶
### Move instructions from CLAUDE.md to skills¶
¶
Your [CLAUDE.md](/en/memory) file is loaded into context at session start. If it contains detailed instructions for specific workflows (like PR reviews or database migrations), those tokens are present even when you're doing unrelated work. [Skills](/en/skills) load on-demand only when invoked, so moving specialized instructions into skills keeps your base context smaller. Aim to keep CLAUDE.md under \~500 lines by including only essentials.¶
¶
### Adjust extended thinking¶
¶
Extended thinking is enabled by default with a budget of 31,999 tokens because it significantly improves performance on complex planning and reasoning tasks. However, thinking tokens are billed as output tokens, so for simpler tasks where deep reasoning isn't needed, you can reduce costs by disabling it in `/config` or lowering the budget (for example, `MAX_THINKING_TOKENS=8000`).¶
¶
### Delegate verbose operations to subagents¶
¶
Running tests, fetching documentation, or processing log files can consume significant context. Delegate these to [subagents](/en/sub-agents#isolate-high-volume-operations) so the verbose output stays in the subagent's context while only a summary returns to your main conversation.¶
¶
### Write specific prompts¶
¶
Vague requests like "improve this codebase" trigger broad scanning. Specific requests like "add input validation to the login function in auth.ts" let Claude work efficiently with minimal file reads.¶
¶
### Work efficiently on complex tasks¶
¶
For longer or more complex work, these habits help avoid wasted tokens from going down the wrong path:¶
¶
* **Use plan mode for complex tasks**: Press Shift+Tab to enter [plan mode](/en/plan-mode) before implementation. Claude explores the codebase and proposes an approach for your approval, preventing expensive re-work when the initial direction is wrong.¶
* **Course-correct early**: If Claude starts heading the wrong direction, press Escape to stop immediately. Use `/rewind` or double-tap Escape to restore conversation and code to a previous checkpoint.¶
* **Give verification targets**: Include test cases, paste screenshots, or define expected output in your prompt. When Claude can verify its own work, it catches issues before you need to request fixes.¶
* **Test incrementally**: Write one file, test it, then continue. This catches issues early when they're cheap to fix.¶
¶
## Background token usage¶
¶
Claude Code uses tokens for some background functionality even when idle:¶
¶
* **Conversation summarization**: Background jobs that summarize previous conversations for the `claude --resume` feature¶
* **Command processing**: Some commands like `/cost` may generate requests to check status¶
¶
These background processes consume a small amount of tokens (typically under \$0.04 per session) even without active interaction.¶
¶
## Understanding changes in Claude Code behavior¶
¶
Claude Code regularly receives updates that may change how features work, including cost reporting. Run `claude --version` to check your current version. For specific billing questions, contact Anthropic support through your [Console account](https://platform.claude.com/login). For team deployments, start with a small pilot group to establish usage patterns before wider rollout.¶
¶
¶
---¶
¶
> To find navigation and other pages in this documentation, fetch the llms.txt file at: https://code.claude.com/docs/llms.txt
Unified Diff
--- a/costs.md
+++ b/costs.md
@@ -1,17 +1,19 @@
# Manage costs effectively
-> Learn how to track and optimize token usage and costs when using Claude Code.
+> Track token usage, set team spend limits, and reduce Claude Code costs with context management, model selection, extended thinking settings, and preprocessing hooks.
-Claude Code consumes tokens for each interaction. The average cost is \$6 per developer per day, with daily costs remaining below \$12 for 90% of users.
+Claude Code consumes tokens for each interaction. Costs vary based on codebase size, query complexity, and conversation length. The average cost is \$6 per developer per day, with daily costs remaining below \$12 for 90% of users.
For team usage, Claude Code charges by API token consumption. On average, Claude Code costs \~\$100-200/developer per month with Sonnet 4.5 though there is large variance depending on how many instances users are running and whether they're using it in automation.
+
+This page covers how to [track your costs](#track-your-costs), [manage costs for teams](#managing-costs-for-teams), and [reduce token usage](#reduce-token-usage).
## Track your costs
### Using the `/cost` command
<Note>
- The `/cost` command is not intended for Claude Max and Pro subscribers.
+ The `/cost` command shows API token usage and is intended for API users. Claude Max and Pro subscribers have usage included in their subscription, so `/cost` data isn't relevant for billing purposes. Subscribers can use `/stats` to view usage patterns.
</Note>
The `/cost` command provides detailed token usage statistics for your current session:
@@ -23,19 +25,15 @@
Total code changes: 0 lines added, 0 lines removed
```
-### Additional tracking options
+## Managing costs for teams
-Check [historical usage](https://support.claude.com/en/articles/9534590-cost-and-usage-reporting-in-console) in the Claude Console (requires Admin or Billing role) and set [workspace spend limits](https://support.claude.com/en/articles/9796807-creating-and-managing-workspaces) for the Claude Code workspace (requires Admin role).
+When using Claude API, you can [set workspace spend limits](https://platform.claude.com/docs/en/build-with-claude/workspaces#workspace-limits) on the total Claude Code workspace spend. Admins can [view cost and usage reporting](https://platform.claude.com/docs/en/build-with-claude/workspaces#usage-and-cost-tracking) in the Console.
<Note>
- When you first authenticate Claude Code with your Claude Console account, a workspace called "Claude Code" is automatically created for you. This workspace provides centralized cost tracking and management for all Claude Code usage in your organization. You cannot create API keys for this workspace - it is exclusively for Claude Code authentication and usage.
+ When you first authenticate Claude Code with your Claude Console account, a workspace called "Claude Code" is automatically created for you. This workspace provides centralized cost tracking and management for all Claude Code usage in your organization. You cannot create API keys for this workspace; it is exclusively for Claude Code authentication and usage.
</Note>
-## Managing costs for teams
-
-When using Claude API, you can limit the total Claude Code workspace spend. To configure, [follow these instructions](https://support.claude.com/en/articles/9796807-creating-and-managing-workspaces). Admins can view cost and usage reporting by [following these instructions](https://support.claude.com/en/articles/9534590-cost-and-usage-reporting-in-console).
-
-On Bedrock and Vertex, Claude Code does not send metrics from your cloud. In order to get cost metrics, several large enterprises reported using [LiteLLM](/en/third-party-integrations#litellm), which is an open-source tool that helps companies [track spend by key](https://docs.litellm.ai/docs/proxy/virtual_keys#tracking-spend). This project is unaffiliated with Anthropic and we have not audited its security.
+On Bedrock, Vertex, and Foundry, Claude Code does not send metrics from your cloud. To get cost metrics, several large enterprises reported using [LiteLLM](/en/llm-gateway#litellm-configuration), which is an open-source tool that helps companies [track spend by key](https://docs.litellm.ai/docs/proxy/virtual_keys#tracking-spend). This project is unaffiliated with Anthropic and we have not audited its security.
### Rate limit recommendations
@@ -60,33 +58,111 @@
## Reduce token usage
-* **Compact conversations:**
+Token costs scale with context size: the more context Claude processes, the more tokens you use. Claude Code automatically optimizes costs through prompt caching (which reduces costs for repeated content like system prompts) and auto-compaction (which summarizes conversation history when approaching context limits).
- * Claude uses auto-compact by default when context reaches approximately 95% capacity. To trigger compaction earlier, set [`CLAUDE_AUTOCOMPACT_PCT_OVERRIDE`](/en/settings#environment-variables) to a lower percentage (for example, `50`)
- * Toggle auto-compact: Run `/config` and navigate to "Auto-compact enabled"
- * Use `/compact` manually when context gets large
- * Add custom instructions: `/compact Focus on code samples and API usage`
- * Customize compaction by adding to CLAUDE.md:
+The following strategies help you keep context small and reduce per-message costs.
- ```markdown theme={null}
- # Summary instructions
+### Manage context proactively
- When you are using compact, please focus on test output and code changes
+Use `/cost` to check your current token usage, or [configure your status line](/en/statusline#context-window-usage) to display it continuously.
+
+* **Clear between tasks**: Use `/clear` to start fresh when switching to unrelated work. Stale context wastes tokens on every subsequent message. Use `/rename` before clearing so you can easily find the session later, then `/resume` to return to it.
+* **Add custom compaction instructions**: `/compact Focus on code samples and API usage` tells Claude what to preserve during summarization.
+
+You can also customize compaction behavior in your CLAUDE.md:
+
+```markdown theme={null}
+# Compact instructions
+
+When you are using compact, please focus on test output and code changes
+```
+
+### Choose the right model
+
+Sonnet handles most coding tasks well and costs less than Opus. Reserve Opus for complex architectural decisions or multi-step reasoning. Use `/model` to switch models mid-session, or set a default in `/config`. For simple subagent tasks, specify `model: haiku` in your [subagent configuration](/en/sub-agents#choose-a-model).
+
+### Reduce MCP server overhead
+
+Each MCP server adds tool definitions to your context, even when idle. Run `/context` to see what's consuming space.
+
+* **Prefer CLI tools when available**: Tools like `gh`, `aws`, `gcloud`, and `sentry-cli` are more context-efficient than MCP servers because they don't add persistent tool definitions. Claude can run CLI commands directly without the overhead.
+* **Disable unused servers**: Run `/mcp` to see configured servers and disable any you're not actively using.
+* **Tool search is automatic**: When MCP tool descriptions exceed 10% of your context window, Claude Code automatically defers them and loads tools on-demand via [tool search](/en/mcp#scale-with-mcp-tool-search). Since deferred tools only enter context when actually used, a lower threshold means fewer idle tool definitions consuming space. Set a lower threshold with `ENABLE_TOOL_SEARCH=auto:<N>` (for example, `auto:5` triggers when tools exceed 5% of your context window).
+
+### Offload processing to hooks and skills
+
+Custom [hooks](/en/hooks) can preprocess data before Claude sees it. Instead of Claude reading a 10,000-line log file to find errors, a hook can grep for `ERROR` and return only matching lines, reducing context from tens of thousands of tokens to hundreds.
+
+A [skill](/en/skills) can give Claude domain knowledge so it doesn't have to explore. For example, a "codebase-overview" skill could describe your project's architecture, key directories, and naming conventions. When Claude invokes the skill, it gets this context immediately instead of spending tokens reading multiple files to understand the structure.
+
+For example, this PreToolUse hook filters test output to show only failures:
+
+<Tabs>
+ <Tab title="settings.json">
+ Add this to your [settings.json](/en/settings#settings-files) to run the hook before every Bash command:
+
+ ```json theme={null}
+ {
+ "hooks": {
+ "PreToolUse": [
+ {
+ "matcher": "Bash",
+ "hooks": [
+ {
+ "type": "command",
+ "command": "~/.claude/hooks/filter-test-output.sh"
+ }
+ ]
+ }
+ ]
+ }
+ }
```
+ </Tab>
-* **Write specific queries:** Avoid vague requests that trigger unnecessary scanning
+ <Tab title="filter-test-output.sh">
+ The hook calls this script, which checks if the command is a test runner and modifies it to show only failures:
-* **Break down complex tasks:** Split large tasks into focused interactions
+ ```bash theme={null}
+ #!/bin/bash
+ input=$(cat)
+ cmd=$(echo "$input" | jq -r '.tool_input.command')
-* **Clear history between tasks:** Use `/clear` to reset context
+ # If running tests, filter to show only failures
+ if [[ "$cmd" =~ ^(npm test|pytest|go test) ]]; then
+ filtered_cmd="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"
+ echo "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"allow\",\"updatedInput\":{\"command\":\"$filtered_cmd\"}}}"
+ else
+ echo "{}"
+ fi
+ ```
+ </Tab>
+</Tabs>
-Costs can vary significantly based on:
+### Move instructions from CLAUDE.md to skills
-* Size of codebase being analyzed
-* Complexity of queries
-* Number of files being searched or modified
-* Length of conversation history
-* Frequency of compacting conversations
+Your [CLAUDE.md](/en/memory) file is loaded into context at session start. If it contains detailed instructions for specific workflows (like PR reviews or database migrations), those tokens are present even when you're doing unrelated work. [Skills](/en/skills) load on-demand only when invoked, so moving specialized instructions into skills keeps your base context smaller. Aim to keep CLAUDE.md under \~500 lines by including only essentials.
+
+### Adjust extended thinking
+
+Extended thinking is enabled by default with a budget of 31,999 tokens because it significantly improves performance on complex planning and reasoning tasks. However, thinking tokens are billed as output tokens, so for simpler tasks where deep reasoning isn't needed, you can reduce costs by disabling it in `/config` or lowering the budget (for example, `MAX_THINKING_TOKENS=8000`).
+
+### Delegate verbose operations to subagents
+
+Running tests, fetching documentation, or processing log files can consume significant context. Delegate these to [subagents](/en/sub-agents#isolate-high-volume-operations) so the verbose output stays in the subagent's context while only a summary returns to your main conversation.
+
+### Write specific prompts
+
+Vague requests like "improve this codebase" trigger broad scanning. Specific requests like "add input validation to the login function in auth.ts" let Claude work efficiently with minimal file reads.
+
+### Work efficiently on complex tasks
+
+For longer or more complex work, these habits help avoid wasted tokens from going down the wrong path:
+
+* **Use plan mode for complex tasks**: Press Shift+Tab to enter [plan mode](/en/plan-mode) before implementation. Claude explores the codebase and proposes an approach for your approval, preventing expensive re-work when the initial direction is wrong.
+* **Course-correct early**: If Claude starts heading the wrong direction, press Escape to stop immediately. Use `/rewind` or double-tap Escape to restore conversation and code to a previous checkpoint.
+* **Give verification targets**: Include test cases, paste screenshots, or define expected output in your prompt. When Claude can verify its own work, it catches issues before you need to request fixes.
+* **Test incrementally**: Write one file, test it, then continue. This catches issues early when they're cheap to fix.
## Background token usage
@@ -97,38 +173,9 @@
These background processes consume a small amount of tokens (typically under \$0.04 per session) even without active interaction.
-## Tracking version changes and updates
+## Understanding changes in Claude Code behavior
-### Current version information
-
-To check your current Claude Code version and installation details:
-
-```bash theme={null}
-claude doctor
-```
-
-This command shows your version, installation type, and system information.
-
-### Understanding changes in Claude Code behavior
-
-Claude Code regularly receives updates that may change how features work, including cost reporting:
-
-* **Version tracking**: Use `claude doctor` to see your current version
-* **Behavior changes**: Features like `/cost` may display information differently across versions
-* **Documentation access**: Claude always has access to the latest documentation, which can help explain current feature behavior
-
-### When cost reporting changes
-
-If you notice changes in how costs are displayed (such as the `/cost` command showing different information):
-
-1. **Verify your version**: Run `claude doctor` to confirm your current version
-2. **Consult documentation**: Ask Claude directly about current feature behavior, as it has access to up-to-date documentation
-3. **Contact support**: For specific billing questions, contact Anthropic support through your Console account
-
-<Note>
- For team deployments, we recommend starting with a small pilot group to
- establish usage patterns before wider rollout.
-</Note>
+Claude Code regularly receives updates that may change how features work, including cost reporting. Run `claude --version` to check your current version. For specific billing questions, contact Anthropic support through your [Console account](https://platform.claude.com/login). For team deployments, start with a small pilot group to establish usage patterns before wider rollout.
---