← Back to daily report
+22 lines added
-19 lines removed
---¶
title: Effort¶
url: https://platform.claude.com/docs/en/build-with-claude/effort¶
description: Control how many tokens Claude uses when responding with the effort parameter, trading off between response thoroughness and token efficiency.¶
featureMetadata:¶
status: ga¶
zdr:¶
eligibility: eligible¶
note: Excludes [Covered Models](https://platform.claude.com/docs/en/manage-claude/api-and-data-retention#model-specific-data-retention-requirements).¶
supportedModels:¶
- claude-fable-5-1¶
- claude-mythos-5-1¶
- claude-fable-5¶
- claude-mythos-5¶
- claude-mythos-preview¶
- claude-opus-5-5¶
- claude-opus-5¶
- claude-opus-4-8¶
- claude-opus-4-7¶
- claude-opus-4-6¶
- claude-opus-4-5-20251101¶
- claude-sonnet-5¶
- claude-sonnet-4-6¶
supportedPlatforms:¶
Claude API: ga¶
Claude Platform on AWS: ga¶
Amazon Bedrock: ga¶
Google Cloud: ga¶
Microsoft Foundry: ga¶
---¶
¶
The effort parameter lets you control how many tokens Claude spends when responding to requests. You can trade off between response thoroughness and token efficiency with a single model. The top-level effort parameter is available on all supported models with no beta header required. [Per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) is in beta.¶
¶
<Tip>¶
To learn how effort interacts with thinking and which control to reach for, see [Thinking and effort](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-effort). Where adaptive thinking is available, effort is the recommended way to control thinking depth.¶
</Tip>¶
¶
## Set the effort level¶
¶
Set `output_config.effort` on the request. The following example runs one request at `medium` effort and prints the response text.¶
¶
<CodeGroup>¶
```bash cURL¶
curl https://api.anthropic.com/v1/messages \¶
-H "x-api-key: $ANTHROPIC_API_KEY" \¶
-H "anthropic-version: 2023-06-01" \¶
-H "content-type: application/json" \¶
-d '{¶
"model": "claude-opus-5-5",¶
"max_tokens": 4096,¶
"messages": [{¶
"role": "user",¶
"content": "Analyze the trade-offs between microservices and monolithic architectures"¶
}],¶
"output_config": {¶
"effort": "medium"¶
}¶
}'¶
```¶
¶
```bash CLI¶
ant messages create \¶
--model claude-opus-5-5 \¶
--max-tokens 4096 \¶
--output-config '{effort: medium}' \¶
--message '{role: user, content: "Analyze the trade-offs between microservices and monolithic architectures"}' \¶
--transform 'content.#(type=="text").text' \¶
--raw-output¶
```¶
¶
```python Python¶
client = anthropic.Anthropic()¶
¶
response = client.messages.create(¶
model="claude-opus-5-5",¶
max_tokens=4096,¶
messages=[¶
{¶
"role": "user",¶
"content": "Analyze the trade-offs between microservices and monolithic architectures",¶
}¶
],¶
output_config={"effort": "medium"},¶
)¶
¶
for block in response.content:¶
if block.type == "text":¶
print(block.text)¶
```¶
¶
```typescript TypeScript¶
const client = new Anthropic();¶
¶
const response = await client.messages.create({¶
model: "claude-opus-5-5",¶
max_tokens: 4096,¶
messages: [¶
{¶
role: "user",¶
content: "Analyze the trade-offs between microservices and monolithic architectures"¶
}¶
],¶
output_config: {¶
effort: "medium"¶
}¶
});¶
¶
const textBlock = response.content.find(¶
(block): block is Anthropic.TextBlock => block.type === "text"¶
);¶
console.log(textBlock?.text);¶
```¶
¶
```csharp C#¶
AnthropicClient client = new();¶
¶
var parameters = new MessageCreateParams¶
{¶
Model = Model.ClaudeOpus5_5,¶
MaxTokens = 4096,¶
Messages = [¶
new() {¶
Role = Role.User,¶
Content = "Analyze the trade-offs between microservices and monolithic architectures"¶
}¶
],¶
OutputConfig = new OutputConfig¶
{¶
Effort = Effort.Medium¶
}¶
};¶
¶
var message = await client.Messages.Create(parameters);¶
Console.WriteLine(message);¶
```¶
¶
```go Go¶
client := anthropic.NewClient()¶
¶
response, err := client.Messages.New(context.TODO(), anthropic.MessageNewParams{¶
Model: anthropic.ModelClaudeOpus5_5,¶
MaxTokens: 4096,¶
Messages: []anthropic.MessageParam{¶
anthropic.NewUserMessage(anthropic.NewTextBlock("Analyze the trade-offs between microservices and monolithic architectures")),¶
},¶
OutputConfig: anthropic.OutputConfigParam{¶
Effort: anthropic.OutputConfigEffortMedium,¶
},¶
})¶
if err != nil {¶
log.Fatal(err)¶
}¶
for _, block := range response.Content {¶
if textBlock, ok := block.AsAny().(anthropic.TextBlock); ok {¶
fmt.Println(textBlock.Text)¶
}¶
}¶
```¶
¶
```java Java¶
import com.anthropic.models.messages.OutputConfig;¶
¶
void main() {¶
AnthropicClient client = AnthropicOkHttpClient.fromEnv();¶
¶
MessageCreateParams params = MessageCreateParams.builder()¶
.model(Model.CLAUDE_OPUS_5_5)¶
.maxTokens(4096L)¶
.addUserMessage("Analyze the trade-offs between microservices and monolithic architectures")¶
.outputConfig(OutputConfig.builder()¶
.effort(OutputConfig.Effort.MEDIUM)¶
.build())¶
.build();¶
¶
Message response = client.messages().create(params);¶
response.content().stream()¶
.flatMap(block -> block.text().stream())¶
.forEach(textBlock -> IO.println(textBlock.text()));¶
}¶
```¶
¶
```php PHP¶
$client = new Client();¶
¶
$message = $client->messages->create(¶
maxTokens: 4096,¶
messages: [¶
['role' => 'user', 'content' => 'Analyze the trade-offs between microservices and monolithic architectures']¶
],¶
model: 'claude-opus-5-5',¶
outputConfig: ['effort' => 'medium'],¶
);¶
¶
foreach ($message->content as $block) {¶
if ($block->type === 'text') {¶
echo $block->text, PHP_EOL;¶
}¶
}¶
```¶
¶
```ruby Ruby¶
client = Anthropic::Client.new¶
¶
message = client.messages.create(¶
model: "claude-opus-5-5",¶
max_tokens: 4096,¶
messages: [¶
{ role: "user", content: "Analyze the trade-offs between microservices and monolithic architectures" }¶
],¶
output_config: {¶
effort: "medium"¶
}¶
)¶
¶
message.content.each do |block|¶
puts block.text if block.type == :text¶
end¶
```¶
</CodeGroup>¶
¶
## How effort works¶
¶
By default, Claude usesMost Claude models default to high effort, spending as many tokens as needed for excellent results; Claude Opus 5.5 defaults to medium. You can raise the effort level to `max` for the absolute highest capability, or lower it to be more conservative with token usage, optimizing for speed and cost while accepting some reduction in capability.¶
¶
<Tip>¶
Setting `effort` to `"high"`the model's default (`"medium"` on Claude Opus 5.5, `"high"` on other models) produces exactly the same behavior as omitting the `effort` parameter entirely.¶
</Tip>¶
¶
The effort parameter affects **all tokens** in the response, including:¶
¶
* Text responses and explanations¶
* Tool calls and function arguments¶
* Thinking (when active)¶
¶
Because effort applies to every output token, it works whether or not thinking is enabled. Lower effort also means fewer and terser tool calls.¶
¶
### Effort levels¶
¶
| Level | Description | Typical use case |¶
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |¶
| `max` | Absolute maximum capability with no constraints on token spending. Available on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, and Claude Sonnet 4.6. | Tasks requiring the deepest possible reasoning and most thorough analysis |¶
| `xhigh` | Extended capability for long-horizon work. Available on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5. | Long-running agentic and coding tasks (over 30 minutes) with token budgets in the millions |¶
| `high` | High capability. Equivalent to not setting the parameter. Spends as many tokens as the task needs for excellent results. The default on every model that supports effort except Claude Opus 5.5. | Complex reasoning, difficult coding problems, agentic tasks |¶
| `medium` | Balanced approach with moderate token savings. The default on Claude Opus 5.5. | Agentic tasks that require a balance of speed, cost, and performance |¶
| `low` | Most efficient. Significant token savings with some capability reduction. | Simpler tasks that need the best speed and lowest costs, such as subagents |¶
¶
Not every model that supports `max` supports `xhigh`.¶
¶
<Note>¶
Effort is a behavioral signal, not a strict token budget. At lower effort levels, Claude still thinks on sufficiently difficult problems, but thinks less than it would at higher effort levels for the same problem.¶
</Note>¶
¶
The per-model recommendations that follow override this table where they differ.¶
¶
### Recommended effort levels for Claude Fable 5.1¶
¶
Claude Fable 5.1 supports all five effort levels. **Start with `high`, the default.** Step up to `xhigh` or `max` for the most capability-sensitive agentic and coding work, and step down to `medium` or `low` for routine or latency-sensitive work once your evals show quality holds. At `high` and above, set a large `max_tokens`. It's a hard limit on total output (thinking plus response text). The same recommendations apply to Claude Mythos 5.1. See [Prompting Claude Fable 5.1](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#consider-all-effort-levels).¶
¶
Claude Fable 5.1 also supports [changing effort mid-conversation](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) with a per-message `output_config`, which preserves the prompt cache.¶
¶
### Recommended effort levels for Claude Fable 5¶
¶
Effort is the primary control for trading off intelligence, latency, and cost on Claude Fable 5. **Start with `high`, the default, for most tasks**, use `xhigh` for the most capability-sensitive workloads, and step down to `medium` or `low` for routine work. Lower effort settings on Claude Fable 5 still perform well and often exceed `xhigh` performance on prior models. At `high` and `xhigh`, set a large `max_tokens`. It's a hard limit on total output (thinking plus response text). See [Cost control](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#cost-control).¶
¶
Reduce effort if a task completes but takes longer than necessary, or if you want a faster, more interactive working style. The same recommendations apply to Claude Mythos 5. For fuller guidance, see [Prompting Claude Fable 5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5).¶
¶
### Recommended effort levels for Claude Opus 5.5¶
¶
Claude Opus 5.5 supports all five effort levels, and `medium` is the default (Claude Opus 5 and earlier Opus models default to `high`, so a request that omits `effort` runs one level lower than it did on Claude Opus 5). Adaptive thinking is always on and can't be turned off, so effort is the primary control for how much the model reasons and what a request costs. Run an effort sweep on your own evals rather than carrying settings over from an earlier model, and set a large `max_tokens` at the higher levels: it's a hard limit on total output (thinking plus response text). Requests that set `thinking: {"type": "disabled"}` return a 400 error at every effort level. Claude Opus 5.5 also supports [changing effort mid-conversation](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) with a per-message `output_config`, which preserves the prompt cache. See [Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5).¶
¶
### Recommended effort levels for Claude Opus 5¶
¶
Claude Opus 5 supports all five effort levels. **Start with `high`, the default**, and adjust based on your evals: step up to `xhigh` for demanding coding and agentic work, or to `max` when a task justifies unconstrained token spending, and use `low` and `medium` liberally as your primary control for token cost and response time wherever your evals show quality holds. If you carried effort settings over from an earlier model, run a fresh effort sweep on your evals rather than reusing them.¶
¶
Effort controls thinking volume, not visible response length: on Claude Opus 5, changing effort does not reliably shorten responses, so [prompt for length](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5#response-length-and-verbosity) instead.¶
¶
The API default is `high`. Set `effort` explicitly to use a different level. The value you pass overrides the default.¶
¶
On Claude Opus 5, thinking cannot be disabled at `xhigh` or `max` effort: requests that set `thinking: {"type": "disabled"}` at those levels return a 400 error. See [Effort with thinking](https://platform.claude.com/docs/en/build-with-claude/effort#effort-with-thinking).¶
¶
When running Claude Opus 5 at `xhigh` or `max` effort, set a large `max_tokens` so the model has room to think and act across subagents and tool calls. Starting at 64k tokens and tuning from there is a reasonable default.¶
¶
Claude Opus 5 also supports [changing effort mid-conversation](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) with a per-message `output_config`, which preserves the prompt cache.¶
¶
### Recommended effort levels for Claude Opus 4.8¶
¶
The guidance for Claude Opus 4.7 also applies to Claude Opus 4.8. **Start with `xhigh` for coding and agentic use cases**, use `high` for most other intelligence-sensitive workloads, and step down to `medium` or `low` only when you've measured that the lower level holds quality on your evals.¶
¶
The API default is `high`. Set `effort` explicitly to use a different level. The value you pass overrides the default.¶
¶
When running Claude Opus 4.8 at `xhigh` or `max` effort, set a large `max_tokens` so the model has room to think and act across subagents and tool calls. Starting at 64k tokens and tuning from there is a reasonable default.¶
¶
### Recommended effort levels for Claude Opus 4.7¶
¶
**Start with `xhigh` for coding and agentic use cases**, and use `high` as the minimum for most intelligence-sensitive workloads. Step down to `medium` for cost-sensitive workloads, or up to `max` only when your evals show measurable headroom at `xhigh`.¶
¶
The API default is `high`. To use `xhigh`, set `effort` explicitly. The value you pass overrides the default.¶
¶
| Effort | Guidance for Claude Opus 4.7 |¶
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |¶
| `low` | Efficient, but best for short, scoped tasks. Pair `low` with explicit checklists if your task has multiple sections. |¶
| `medium` | The drop-in for the average workflow where you want good results while reducing costs. |¶
| `high` | Advanced use cases that still need a balance of intelligence and token consumption. This is often the best balance of quality and token efficiency. |¶
| `xhigh` | The recommended starting point for coding and agentic work, and for exploratory tasks such as repeated tool calling, detailed web search, and knowledge-base search. Expect meaningfully higher token usage than `high`. |¶
| `max` | Reserve for frontier problems. On most workloads `max` adds significant cost for relatively small quality gains, and on some structured-output or less intelligence-sensitive tasks it can lead to overthinking. |¶
¶
Claude Opus 4.7 also respects effort levels more strictly than Claude Opus 4.6, especially at `low` and `medium`. At lower effort levels, the model scopes its work to what was asked rather than doing more than requested. If you observe shallow reasoning on complex problems with Claude Opus 4.7, raise effort rather than prompting around it. If you must keep effort low for latency, add targeted guidance like "This task involves multistep reasoning. Think carefully before responding."¶
¶
When running Claude Opus 4.7 at `xhigh` or `max` effort, set a large `max_tokens` so the model has room to think and act across subagents and tool calls. Starting at 64k tokens and tuning from there is a reasonable default.¶
¶
### Recommended effort levels for Claude Sonnet 5¶
¶
Claude Sonnet 5 defaults to `high` effort on the Claude API and Claude Code.¶
¶
* **High effort (default):** Suitable for complex reasoning, coding, and agentic tasks where quality matters more than speed or cost.¶
* **Xhigh effort:** For the hardest coding and agentic tasks. See [Prompting Claude Sonnet 5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5#calibrating-effort-and-thinking-depth).¶
* **Medium effort:** Cost-saving step-down from the default. Comparable to Claude Sonnet 4.6 at high effort.¶
* **Low effort:** For high-volume or latency-sensitive workloads. Suitable for chat and non-coding use cases where faster turnaround is prioritized.¶
* **Max effort:** For tasks requiring the absolute highest capability with no constraints on token spending.¶
¶
### Recommended effort levels for Claude Sonnet 4.6¶
¶
Sonnet 4.6 defaults to `high` effort. Explicitly set effort when using Sonnet 4.6 to avoid unexpected latency:¶
¶
* **Medium effort** (recommended default): Best balance of speed, cost, and performance for most applications. Suitable for agentic coding, tool-heavy workflows, and code generation.¶
* **Low effort:** For high-volume or latency-sensitive workloads. Suitable for chat and non-coding use cases where faster turnaround is prioritized.¶
* **High effort:** For complex reasoning and tasks where quality matters more than speed or cost.¶
* **Max effort:** For tasks requiring the absolute highest capability with no constraints on token spending.¶
¶
## Effort with tool use¶
¶
When using tools, the effort parameter affects both the explanations around tool calls and the tool calls themselves. Lower effort levels tend to:¶
¶
* Combine multiple operations into fewer tool calls¶
* Make fewer tool calls¶
* Proceed directly to action without preamble¶
* Use terse confirmation messages after completion¶
¶
Higher effort levels may:¶
¶
* Make more tool calls¶
* Explain the plan before taking action¶
* Provide detailed summaries of changes¶
* Include more comprehensive code comments¶
¶
## Effort with thinking¶
¶
The `thinking` parameter controls whether Claude thinks in [thinking blocks](https://platform.claude.com/docs/en/build-with-claude/thinking) before answering; the `effort` parameter controls how much work Claude puts into the whole response, which in adaptive mode includes how often and how deeply it thinks. Don't pass `adaptive` as an `effort` value: `adaptive` is a thinking mode, not an effort level.¶
¶
At higher effort levels, Claude thinks more readily and at greater length. In a tool-use loop, follow-up requests that only process tool results can still skip thinking at any level. At lower levels, Claude can skip thinking entirely for simpler problems. See [Thinking and effort](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-effort) for full guidance on how the two controls work together.¶
¶
On Claude Opus 4.5, the only extended-thinking-only model that supports effort, it works alongside [`budget_tokens`](https://platform.claude.com/docs/en/build-with-claude/extended-thinking): set the effort level for your task, then set the thinking token budget based on how much reasoning depth the task needs.¶
¶
For per-model thinking availability, see the [per-model configuration table](https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting#supported-models). Effort works with or without thinking. See [How effort works](https://platform.claude.com/docs/en/build-with-claude/effort#how-effort-works).¶
¶
## Change effort mid-conversation¶
¶
You can run later turns of a conversation at a different effort level in two ways. On Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, and Claude Opus 5, use a per-message effort change, which keeps the prompt cache. On other models, set a new top-level value on the next request, which starts the cache over.¶
¶
### Per-message effort (beta)¶
¶
Per-message effort is in beta and requires the [beta header](https://platform.claude.com/docs/en/api/beta-headers) `mid-conversation-output-config-2026-07-01`. Models without per-message effort, including Claude Fable 5, return a 400 error: `output_config.effort requires a model that supports per-turn effort; this model does not`.¶
¶
Add a `role: "system"` message with empty `content` and the new level in `output_config.effort`. The new level takes effect from the next `user` turn and holds until a later message changes it. Everything before that message is unchanged, so the cached prefix still matches.¶
¶
The following example starts at `high`, then drops to `low` for a routine follow-up:¶
¶
<CodeGroup>¶
```bash cURL¶
# Effort-only system message: the new level takes effect from the next user turn.¶
curl https://api.anthropic.com/v1/messages \¶
-H "x-api-key: $ANTHROPIC_API_KEY" \¶
-H "anthropic-version: 2023-06-01" \¶
-H "anthropic-beta: mid-conversation-output-config-2026-07-01" \¶
-H "content-type: application/json" \¶
-d '{¶
"model": "claude-fable-5-1",¶
"max_tokens": 4096,¶
"output_config": {"effort": "high"},¶
"messages": [¶
{"role": "user", "content": "Plan a migration from SQLite to PostgreSQL in three short steps."},¶
{"role": "assistant", "content": "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."},¶
{"role": "system", "content": [], "output_config": {"effort": "low"}},¶
{"role": "user", "content": "Summarize the plan in one sentence."}¶
]¶
}'¶
```¶
¶
```bash CLI¶
ant beta:messages create --beta mid-conversation-output-config-2026-07-01 \¶
--transform 'content.#(type=="text").text' --raw-output <<'YAML'¶
model: claude-fable-5-1¶
max_tokens: 4096¶
output_config:¶
effort: high¶
messages:¶
- role: user¶
content: Plan a migration from SQLite to PostgreSQL in three short steps.¶
- role: assistant¶
content: "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."¶
# Effort-only system message: the new level takes effect from the next user turn.¶
- role: system¶
content: []¶
output_config:¶
effort: low¶
- role: user¶
content: Summarize the plan in one sentence.¶
YAML¶
```¶
¶
```python Python¶
client = anthropic.Anthropic()¶
¶
response = client.beta.messages.create(¶
model="claude-fable-5-1",¶
max_tokens=4096,¶
output_config={"effort": "high"},¶
messages=[¶
{¶
"role": "user",¶
"content": "Plan a migration from SQLite to PostgreSQL in three short steps.",¶
},¶
{¶
"role": "assistant",¶
"content": "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.",¶
},¶
# Effort-only system message: the new level takes effect from the next user turn.¶
{"role": "system", "content": [], "output_config": {"effort": "low"}},¶
{"role": "user", "content": "Summarize the plan in one sentence."},¶
],¶
betas=["mid-conversation-output-config-2026-07-01"],¶
)¶
¶
for block in response.content:¶
if block.type == "text":¶
print(block.text)¶
```¶
¶
```typescript TypeScript¶
const client = new Anthropic();¶
¶
const response = await client.beta.messages.create({¶
model: "claude-fable-5-1",¶
max_tokens: 4096,¶
output_config: { effort: "high" },¶
messages: [¶
{¶
role: "user",¶
content: "Plan a migration from SQLite to PostgreSQL in three short steps."¶
},¶
{¶
role: "assistant",¶
content:¶
"1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."¶
},¶
// Effort-only system message: the new level takes effect from the next user turn.¶
{ role: "system", content: [], output_config: { effort: "low" } },¶
{ role: "user", content: "Summarize the plan in one sentence." }¶
],¶
betas: ["mid-conversation-output-config-2026-07-01"]¶
});¶
¶
for (const block of response.content) {¶
if (block.type === "text") {¶
console.log(block.text);¶
}¶
}¶
```¶
¶
```csharp C#¶
using Anthropic.Models.Beta;¶
using Anthropic.Models.Beta.Messages;¶
¶
AnthropicClient client = new();¶
¶
var response = await client.Beta.Messages.Create(new MessageCreateParams¶
{¶
Model = "claude-fable-5-1",¶
MaxTokens = 4096,¶
OutputConfig = new() { Effort = Effort.High },¶
Messages =¶
[¶
new() { Role = Role.User, Content = "Plan a migration from SQLite to PostgreSQL in three short steps." },¶
new() { Role = Role.Assistant, Content = "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts." },¶
// Effort-only system message: the new level takes effect from the next user turn.¶
new()¶
{¶
Role = Role.System,¶
Content = new([]),¶
OutputConfig = new() { Effort = BetaSystemMessageOutputConfigEffort.Low },¶
},¶
new() { Role = Role.User, Content = "Summarize the plan in one sentence." },¶
],¶
Betas = [AnthropicBeta.MidConversationOutputConfig2026_07_01],¶
});¶
¶
foreach (var block in response.Content)¶
{¶
if (block.TryPickText(out var textBlock))¶
{¶
Console.WriteLine(textBlock.Text);¶
}¶
}¶
```¶
¶
```go Go¶
client := anthropic.NewClient()¶
¶
response, err := client.Beta.Messages.New(context.Background(), anthropic.BetaMessageNewParams{¶
Model: "claude-fable-5-1",¶
MaxTokens: 4096,¶
OutputConfig: anthropic.BetaOutputConfigParam{¶
Effort: anthropic.BetaOutputConfigEffortHigh,¶
},¶
Messages: []anthropic.BetaMessageParam{¶
anthropic.NewBetaUserMessage(anthropic.NewBetaTextBlock("Plan a migration from SQLite to PostgreSQL in three short steps.")),¶
{¶
Role: anthropic.BetaMessageParamRoleAssistant,¶
Content: []anthropic.BetaContentBlockParamUnion{anthropic.NewBetaTextBlock("1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.")},¶
},¶
// Effort-only system message: the new level takes effect from the next user turn.¶
anthropic.NewBetaSystemMessage(anthropic.BetaSystemMessageOutputConfigParam{¶
Effort: anthropic.BetaSystemMessageOutputConfigEffortLow,¶
}),¶
anthropic.NewBetaUserMessage(anthropic.NewBetaTextBlock("Summarize the plan in one sentence.")),¶
},¶
Betas: []anthropic.AnthropicBeta{anthropic.AnthropicBetaMidConversationOutputConfig2026_07_01},¶
})¶
if err != nil {¶
log.Fatal(err)¶
}¶
¶
for _, block := range response.Content {¶
if textBlock, ok := block.AsAny().(anthropic.BetaTextBlock); ok {¶
fmt.Println(textBlock.Text)¶
}¶
}¶
```¶
¶
```java Java¶
import com.anthropic.models.beta.AnthropicBeta;¶
import com.anthropic.models.beta.messages.BetaMessage;¶
import com.anthropic.models.beta.messages.BetaMessageParam;¶
import com.anthropic.models.beta.messages.BetaOutputConfig;¶
import com.anthropic.models.beta.messages.BetaSystemMessageOutputConfig;¶
import com.anthropic.models.beta.messages.MessageCreateParams;¶
¶
void main() {¶
AnthropicClient client = AnthropicOkHttpClient.fromEnv();¶
¶
MessageCreateParams params = MessageCreateParams.builder()¶
.model("claude-fable-5-1")¶
.maxTokens(4096L)¶
.addBeta(AnthropicBeta.MID_CONVERSATION_OUTPUT_CONFIG_2026_07_01)¶
.outputConfig(BetaOutputConfig.builder()¶
.effort(BetaOutputConfig.Effort.HIGH)¶
.build())¶
.addUserMessage("Plan a migration from SQLite to PostgreSQL in three short steps.")¶
.addAssistantMessage("1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.")¶
// Effort-only system message: the new level takes effect from the next user turn.¶
.addMessage(BetaMessageParam.builder()¶
.role(BetaMessageParam.Role.SYSTEM)¶
.contentOfBetaContentBlockParams(List.of())¶
.outputConfig(BetaSystemMessageOutputConfig.builder()¶
.effort(BetaSystemMessageOutputConfig.Effort.LOW)¶
.build())¶
.build())¶
.addUserMessage("Summarize the plan in one sentence.")¶
.build();¶
¶
BetaMessage response = client.beta().messages().create(params);¶
response.content().stream()¶
.flatMap(block -> block.text().stream())¶
.forEach(textBlock -> IO.println(textBlock.text()));¶
}¶
```¶
¶
```php PHP¶
use Anthropic\Beta\AnthropicBeta;¶
use Anthropic\Beta\Messages\BetaMessageParam;¶
use Anthropic\Beta\Messages\BetaOutputConfig;¶
use Anthropic\Beta\Messages\BetaSystemMessageOutputConfig;¶
use Anthropic\Client;¶
¶
$client = new Client();¶
¶
$response = $client->beta->messages->create(¶
model: 'claude-fable-5-1',¶
maxTokens: 4096,¶
outputConfig: BetaOutputConfig::with(effort: 'high'),¶
messages: [¶
BetaMessageParam::with(role: 'user', content: 'Plan a migration from SQLite to PostgreSQL in three short steps.'),¶
BetaMessageParam::with(role: 'assistant', content: '1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.'),¶
// Effort-only system message: the new level takes effect from the next user turn.¶
BetaMessageParam::with(¶
role: 'system',¶
content: [],¶
outputConfig: BetaSystemMessageOutputConfig::with(effort: 'low'),¶
),¶
BetaMessageParam::with(role: 'user', content: 'Summarize the plan in one sentence.'),¶
],¶
betas: [AnthropicBeta::MID_CONVERSATION_OUTPUT_CONFIG_2026_07_01],¶
);¶
¶
foreach ($response->content as $block) {¶
if ($block->type === 'text') {¶
echo $block->text, PHP_EOL;¶
}¶
}¶
```¶
¶
```ruby Ruby¶
client = Anthropic::Client.new¶
¶
response = client.beta.messages.create(¶
model: "claude-fable-5-1",¶
max_tokens: 4096,¶
output_config: {effort: :high},¶
messages: [¶
{role: "user", content: "Plan a migration from SQLite to PostgreSQL in three short steps."},¶
{role: "assistant", content: "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."},¶
# Effort-only system message: the new level takes effect from the next user turn.¶
{role: "system", content: [], output_config: {effort: :low}},¶
{role: "user", content: "Summarize the plan in one sentence."}¶
],¶
betas: [Anthropic::AnthropicBeta::MID_CONVERSATION_OUTPUT_CONFIG_2026_07_01]¶
)¶
¶
response.content.each do |block|¶
puts block.text if block.type == :text¶
end¶
```¶
</CodeGroup>¶
¶
An effort-only system message carries no text, so the [placement rules for mid-conversation system messages](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#limitations) don't apply. It can appear anywhere in `messages`, including as the first entry or between an `assistant` turn and the next `user` turn. Values are the level names (`low`, `medium`, `high`, `xhigh`, and `max`).¶
¶
On Claude Fable 5.1, prefer this form over changing the top-level value between requests. A top-level change restarts the cache and also steers the model less reliably: its earlier replies were written at the previous level, and it tends to stay consistent with them.¶
¶
### Top-level effort on the next request¶
¶
The top-level `output_config.effort` applies to the whole request. To run a later part of a conversation at a different level, set the new value on the next request. Because top-level effort shapes the rendered prompt, changing it between requests doesn't preserve cached prefixes from earlier turns. If you rely on [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) across a long session and your model doesn't support per-message effort, pick an effort level at the start and keep it constant.¶
¶
## Best practices¶
¶
1. **Set effort explicitly:** The API defaults to `high` (`medium` on Claude Opus 5.5), but the right starting point depends on your model and workload.¶
2. **Use low for speed-sensitive or simple tasks:** When latency matters or tasks are straightforward, low effort can significantly reduce response times and costs.¶
3. **Test your use case:** The impact of effort levels varies by task type. Evaluate performance on your specific use cases before deploying.¶
4. **Consider dynamic effort:** Adjust effort based on task complexity. Simple queries may warrant low effort while agentic coding and complex reasoning benefit from high effort. See the next item before varying it within one conversation.¶
5. **Hold top-level effort constant within cached conversations:** Changing the top-level effort value between requests invalidates [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching), so vary it across workloads rather than within a conversation that relies on cache hits. On models that support it, use a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) instead, which preserves the cache. See [Thinking and prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching).¶
¶
## Next steps¶
¶
<CardGroup>¶
<Card title="Task budgets" icon="gauge" href="https://platform.claude.com/docs/en/build-with-claude/task-budgets">¶
Give Claude an advisory token budget for the full agentic loop to help the model self-regulate on long agentic tasks.¶
</Card>¶
¶
<Card title="Steering thinking" icon="compass" href="https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost">¶
Understand adaptive thinking, where Claude decides when and how much to think, and steer it with effort and prompting.¶
</Card>¶
¶
<Card title="Thinking" icon="brain" href="https://platform.claude.com/docs/en/build-with-claude/thinking">¶
Understand how thinking works, when Claude thinks by default, and how thinking interacts with effort.¶
</Card>¶
</CardGroup>¶
Unified Diff
--- a/build-with-claude/effort.md
+++ b/build-with-claude/effort.md
@@ -13,6 +13,7 @@
- claude-fable-5
- claude-mythos-5
- claude-mythos-preview
+ - claude-opus-5-5
- claude-opus-5
- claude-opus-4-8
- claude-opus-4-7
@@ -45,7 +46,7 @@
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
- "model": "claude-opus-5",
+ "model": "claude-opus-5-5",
"max_tokens": 4096,
"messages": [{
"role": "user",
@@ -59,7 +60,7 @@
```bash CLI
ant messages create \
- --model claude-opus-5 \
+ --model claude-opus-5-5 \
--max-tokens 4096 \
--output-config '{effort: medium}' \
--message '{role: user, content: "Analyze the trade-offs between microservices and monolithic architectures"}' \
@@ -71,7 +72,7 @@
client = anthropic.Anthropic()
response = client.messages.create(
- model="claude-opus-5",
+ model="claude-opus-5-5",
max_tokens=4096,
messages=[
{
@@ -91,7 +92,7 @@
const client = new Anthropic();
const response = await client.messages.create({
- model: "claude-opus-5",
+ model: "claude-opus-5-5",
max_tokens: 4096,
messages: [
{
@@ -115,7 +116,7 @@
var parameters = new MessageCreateParams
{
- Model = Model.ClaudeOpus5,
+ Model = Model.ClaudeOpus5_5,
MaxTokens = 4096,
Messages = [
new() {
@@ -137,7 +138,7 @@
client := anthropic.NewClient()
response, err := client.Messages.New(context.TODO(), anthropic.MessageNewParams{
- Model: anthropic.ModelClaudeOpus5,
+ Model: anthropic.ModelClaudeOpus5_5,
MaxTokens: 4096,
Messages: []anthropic.MessageParam{
anthropic.NewUserMessage(anthropic.NewTextBlock("Analyze the trade-offs between microservices and monolithic architectures")),
@@ -163,7 +164,7 @@
AnthropicClient client = AnthropicOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
- .model(Model.CLAUDE_OPUS_5)
+ .model(Model.CLAUDE_OPUS_5_5)
.maxTokens(4096L)
.addUserMessage("Analyze the trade-offs between microservices and monolithic architectures")
.outputConfig(OutputConfig.builder()
@@ -186,7 +187,7 @@
messages: [
['role' => 'user', 'content' => 'Analyze the trade-offs between microservices and monolithic architectures']
],
- model: 'claude-opus-5',
+ model: 'claude-opus-5-5',
outputConfig: ['effort' => 'medium'],
);
@@ -201,7 +202,7 @@
client = Anthropic::Client.new
message = client.messages.create(
- model: "claude-opus-5",
+ model: "claude-opus-5-5",
max_tokens: 4096,
messages: [
{ role: "user", content: "Analyze the trade-offs between microservices and monolithic architectures" }
@@ -219,10 +220,10 @@
## How effort works
-By default, Claude uses high effort, spending as many tokens as needed for excellent results. You can raise the effort level to `max` for the absolute highest capability, or lower it to be more conservative with token usage, optimizing for speed and cost while accepting some reduction in capability.
+Most Claude models default to high effort, spending as many tokens as needed for excellent results; Claude Opus 5.5 defaults to medium. You can raise the effort level to `max` for the absolute highest capability, or lower it to be more conservative with token usage, optimizing for speed and cost while accepting some reduction in capability.
<Tip>
- Setting `effort` to `"high"` produces exactly the same behavior as omitting the `effort` parameter entirely.
+ Setting `effort` to the model's default (`"medium"` on Claude Opus 5.5, `"high"` on other models) produces exactly the same behavior as omitting the `effort` parameter entirely.
</Tip>
The effort parameter affects **all tokens** in the response, including:
@@ -235,13 +236,13 @@
### Effort levels
-| Level | Description | Typical use case |
-| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
-| `max` | Absolute maximum capability with no constraints on token spending. Available on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, and Claude Sonnet 4.6. | Tasks requiring the deepest possible reasoning and most thorough analysis |
-| `xhigh` | Extended capability for long-horizon work. Available on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5. | Long-running agentic and coding tasks (over 30 minutes) with token budgets in the millions |
-| `high` | High capability. Equivalent to not setting the parameter. | Complex reasoning, difficult coding problems, agentic tasks |
-| `medium` | Balanced approach with moderate token savings. | Agentic tasks that require a balance of speed, cost, and performance |
-| `low` | Most efficient. Significant token savings with some capability reduction. | Simpler tasks that need the best speed and lowest costs, such as subagents |
+| Level | Description | Typical use case |
+| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
+| `max` | Absolute maximum capability with no constraints on token spending. Available on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, and Claude Sonnet 4.6. | Tasks requiring the deepest possible reasoning and most thorough analysis |
+| `xhigh` | Extended capability for long-horizon work. Available on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5. | Long-running agentic and coding tasks (over 30 minutes) with token budgets in the millions |
+| `high` | Spends as many tokens as the task needs for excellent results. The default on every model that supports effort except Claude Opus 5.5. | Complex reasoning, difficult coding problems, agentic tasks |
+| `medium` | Balanced approach with moderate token savings. The default on Claude Opus 5.5. | Agentic tasks that require a balance of speed, cost, and performance |
+| `low` | Most efficient. Significant token savings with some capability reduction. | Simpler tasks that need the best speed and lowest costs, such as subagents |
Not every model that supports `max` supports `xhigh`.
@@ -262,6 +263,10 @@
Effort is the primary control for trading off intelligence, latency, and cost on Claude Fable 5. **Start with `high`, the default, for most tasks**, use `xhigh` for the most capability-sensitive workloads, and step down to `medium` or `low` for routine work. Lower effort settings on Claude Fable 5 still perform well and often exceed `xhigh` performance on prior models. At `high` and `xhigh`, set a large `max_tokens`. It's a hard limit on total output (thinking plus response text). See [Cost control](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#cost-control).
Reduce effort if a task completes but takes longer than necessary, or if you want a faster, more interactive working style. The same recommendations apply to Claude Mythos 5. For fuller guidance, see [Prompting Claude Fable 5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5).
+
+### Recommended effort levels for Claude Opus 5.5
+
+Claude Opus 5.5 supports all five effort levels, and `medium` is the default (Claude Opus 5 and earlier Opus models default to `high`, so a request that omits `effort` runs one level lower than it did on Claude Opus 5). Adaptive thinking is always on and can't be turned off, so effort is the primary control for how much the model reasons and what a request costs. Run an effort sweep on your own evals rather than carrying settings over from an earlier model, and set a large `max_tokens` at the higher levels: it's a hard limit on total output (thinking plus response text). Requests that set `thinking: {"type": "disabled"}` return a 400 error at every effort level. Claude Opus 5.5 also supports [changing effort mid-conversation](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) with a per-message `output_config`, which preserves the prompt cache. See [Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5).
### Recommended effort levels for Claude Opus 5
@@ -350,7 +355,7 @@
## Change effort mid-conversation
-You can run later turns of a conversation at a different effort level in two ways. On Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5, use a per-message effort change, which keeps the prompt cache. On other models, set a new top-level value on the next request, which starts the cache over.
+You can run later turns of a conversation at a different effort level in two ways. On Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, and Claude Opus 5, use a per-message effort change, which keeps the prompt cache. On other models, set a new top-level value on the next request, which starts the cache over.
### Per-message effort (beta)
@@ -637,7 +642,7 @@
## Best practices
-1. **Set effort explicitly:** The API defaults to `high`, but the right starting point depends on your model and workload.
+1. **Set effort explicitly:** The API defaults to `high` (`medium` on Claude Opus 5.5), but the right starting point depends on your model and workload.
2. **Use low for speed-sensitive or simple tasks:** When latency matters or tasks are straightforward, low effort can significantly reduce response times and costs.
3. **Test your use case:** The impact of effort levels varies by task type. Evaluate performance on your specific use cases before deploying.
4. **Consider dynamic effort:** Adjust effort based on task complexity. Simple queries may warrant low effort while agentic coding and complex reasoning benefit from high effort. See the next item before varying it within one conversation.