// five levels, one default
output_config.effort takes five values. The API default is high.
effort to "high" produces exactly the same behavior as omitting the effort parameter entirely.So there is no opt-in moment to remember. Every call you have made ran at high, and the only way to be anywhere else is to say so.
Top-level effort is generally available — no beta header — on Opus 5, 4.8, 4.7, 4.6, Sonnet 5 and 4.6, Fable and Mythos 5 / 5.1, and Opus 4.5. One wrinkle worth knowing before you sweep the range: not every model that supports max supports xhigh.
It moves everything it writes
The common mental model is that effort is a thinking dial. It is broader than that, which is what makes it useful for agents rather than just for reasoning.
Because it applies to every output token, it works whether or not thinking is enabled. Lowering it changes how the agent behaves, not only how long it deliberates:
A signal, not a cap
This is the line that keeps the parameter honest, and the one that separates it from anything with the word budget on it:
So low is not a ceiling you can plan capacity against. Hand a genuinely hard problem to a low-effort request and it will still be worked through — just less thoroughly than the same request one rung up. If what you need is a hard limit on how far an agent may go, that is a different control: see task budgets, which counts what the agent writes and what your tools hand back.
It is not the length dial
The instinct after reading the level table is to turn effort down when a model is being verbose. Wrong lever:
If you want a shorter answer, ask for a shorter answer. Effort changes how hard the model works on the response; the response's shape and length are things you specify.
Two things that bite later
Changing it mid-conversation throws away your cache. Top-level effort shapes the rendered prompt, so a new value on the next request does not match the cached prefix from earlier turns.
On models that support it (Fable 5.1, Mythos 5.1, Opus 5) there is a per-message form that preserves the cache — a role: "system" message with empty content carrying the new level, behind the beta header mid-conversation-output-config-2026-07-01. The new level takes effect from the next user turn; everything before it is unchanged, so the prefix still matches.
At the top of the range, thinking is not optional. On Claude Opus 5, a request that sets thinking: {"type": "disabled"} at xhigh or max returns a 400. Also set a large max_tokens up there — it is a hard limit on total output, thinking plus response text, and 64k is a reasonable starting point.
And carry nothing over blindly: if you inherited effort settings from an earlier model, run a fresh sweep on your evals rather than reusing them.
Related: stop sending every step to your best model is the other half of this decision — that one picks a different model, this one keeps the same model and works it less hard. Why the reply costs five times the prompt is the bill underneath both.
Source: Claude Platform docs — Effort. Verified 2026-09-12.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
