toolcall() ← all concepts

// concept · workflows

You are already on high

There is a setting on every request you have ever sent that you have almost certainly never touched. Its default is not off, and it is not the bottom of the range — it is the middle, with two levels above it.

// five levels, one default

output_config.effort takes five values. The API default is high.

max absolute maximum capability, no constraint on token spending xhigh extended capability for long-horizon agentic and coding work high the default — complex reasoning, hard coding, agentic tasks medium balanced, moderate token savings low most efficient, significant savings with some capability reduction
Setting effort to "high" produces exactly the same behavior as omitting the effort parameter entirely.

So there is no opt-in moment to remember. Every call you have made ran at high, and the only way to be anywhere else is to say so.

response = client.messages.create( model="claude-opus-5", max_tokens=4096, messages=[...], output_config={"effort": "medium"}, )

Top-level effort is generally available — no beta header — on Opus 5, 4.8, 4.7, 4.6, Sonnet 5 and 4.6, Fable and Mythos 5 / 5.1, and Opus 4.5. One wrinkle worth knowing before you sweep the range: not every model that supports max supports xhigh.

It moves everything it writes

The common mental model is that effort is a thinking dial. It is broader than that, which is what makes it useful for agents rather than just for reasoning.

The effort parameter affects all tokens in the response, including: text responses and explanations · tool calls and function arguments · thinking (when active).

Because it applies to every output token, it works whether or not thinking is enabled. Lowering it changes how the agent behaves, not only how long it deliberates:

lower effort combines operations into fewer tool calls makes fewer tool calls proceeds straight to action without preamble terse confirmation after completion higher effort more tool calls explains the plan before acting detailed summaries of what changed

A signal, not a cap

This is the line that keeps the parameter honest, and the one that separates it from anything with the word budget on it:

Effort is a behavioral signal, not a strict token budget. At lower effort levels, Claude still thinks on sufficiently difficult problems, but thinks less than it would at higher effort levels for the same problem.

So low is not a ceiling you can plan capacity against. Hand a genuinely hard problem to a low-effort request and it will still be worked through — just less thoroughly than the same request one rung up. If what you need is a hard limit on how far an agent may go, that is a different control: see task budgets, which counts what the agent writes and what your tools hand back.

It is not the length dial

The instinct after reading the level table is to turn effort down when a model is being verbose. Wrong lever:

Effort controls thinking volume, not visible response length: on Claude Opus 5, changing effort does not reliably shorten responses, so prompt for length instead.

If you want a shorter answer, ask for a shorter answer. Effort changes how hard the model works on the response; the response's shape and length are things you specify.

effort: "low" # thinks less, may still write at length "Answer in two sentences." # this is the length control

Two things that bite later

Changing it mid-conversation throws away your cache. Top-level effort shapes the rendered prompt, so a new value on the next request does not match the cached prefix from earlier turns.

Changing the top-level effort value between requests invalidates prompt caching, so vary it across workloads rather than within a conversation that relies on cache hits.

On models that support it (Fable 5.1, Mythos 5.1, Opus 5) there is a per-message form that preserves the cache — a role: "system" message with empty content carrying the new level, behind the beta header mid-conversation-output-config-2026-07-01. The new level takes effect from the next user turn; everything before it is unchanged, so the prefix still matches.

{"role": "system", "content": [], "output_config": {"effort": "low"}}

At the top of the range, thinking is not optional. On Claude Opus 5, a request that sets thinking: {"type": "disabled"} at xhigh or max returns a 400. Also set a large max_tokens up there — it is a hard limit on total output, thinking plus response text, and 64k is a reasonable starting point.

And carry nothing over blindly: if you inherited effort settings from an earlier model, run a fresh sweep on your evals rather than reusing them.

Related: stop sending every step to your best model is the other half of this decision — that one picks a different model, this one keeps the same model and works it less hard. Why the reply costs five times the prompt is the bill underneath both.

Source: Claude Platform docs — Effort. Verified 2026-09-12.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click