toolcall() ← all concepts

// concept · agentic

Your token budget doesn't count your conversation

You hand your agent a 64,000-token budget and picture 64,000 tokens of conversation. That is not what is being counted — and the gap explains why budgets run out much later than people expect.

// what everyone assumes

The Messages API is stateless. Every request you send carries the whole conversation again — system prompt, every prior turn, every tool result from earlier laps. So when a parameter called a budget appears, the natural reading is that it caps that pile.

# the intuitive (wrong) model budget 64,000 = how big the conversation may get

It isn't. The budget is scoped to the turn, not to the transcript.

What it actually counts

Two things, and only on the turn in progress: what Claude generates, and the tool results it reads back. The history you resend rides along free.

counted tokens Claude writes this turn counted tool results returned to it this turn free the entire history you resent to get here

Anthropic's own worked example makes the size of the gap concrete: a three-turn loop that puts roughly 20,820 tokens of payload on the wire counts 19,000 against the budget. The difference is the resend.

This is the exact inverse of the thing that makes agent runs expensive. A twenty-call agent bills like two hundred precisely because the context is resent and re-billed every lap — and a task budget deliberately does not count that regrowth. Your bill and your budget are measuring different things.

A budget the model can see

This is what separates it from the cap you already have. max_tokens is an enforced per-response ceiling the model is not aware of. Hit it and the response stops mid-thought and comes back with stop_reason: "max_tokens". Nothing warned it; nothing wrapped up.

max_tokens hard, invisible → truncated mid-sentence task_budget soft, visible → paces, then finishes

With a task budget the server injects a countdown marker Claude sees while generating, so it can prioritise and close cleanly instead of being guillotined. That difference — a limit the model can plan around versus one it walks into — is the whole reason the parameter exists. It is the API-level version of the point in nothing is telling your agent to stop.

# output_config, beta header task-budgets-2026-03-13 output_config: { task_budget: { type: "tokens", total: 64000 } }

The trap: leave remaining alone

remaining defaults to total, and the server tracks the countdown itself. The temptation is to be helpful and decrement it yourself. Don't — not in an ordinary loop that resends history.

From the docs: if you also decrement remaining while resending full history, the model sees an under-reported budget and the countdown drops faster than it should, causing Claude to wrap up earlier than the budget actually allows.

The failure is quiet in the worst way: nothing errors, the agent just finishes early and you conclude the budget was too small. Pass remaining only when you compact or rewrite history, so the server can no longer derive prior spend on its own.

The fine print

Four things that decide whether this is available to you at all:

beta header task-budgets-2026-03-13 advisory it paces spend — it does not prevent it minimum total must be ≥ 20,000 tokens (400 below that) vendor Anthropic-specific, not an industry parameter

Model support is narrower than people expect, and it is worth checking before you plan around it:

supported Opus 5 · Fable 5 · Mythos 5 · Opus 4.8 · Opus 4.7 not supported Sonnet 5 · Opus 4.6 · Sonnet 4.6 · Haiku 4.5

Task budgets are also not available on the Claude Code or Cowork surfaces. And do not confuse this with Managed Agents session budgets, which are a different thing entirely: hard, dollar-denominated, platform-enforced caps on a session. A task budget is advisory and token-denominated. Say "paces", never "prevents".

Sources: Claude API — Task budgets (beta) · Claude API — Handling stop reasons · Claude API — Managed Agents session budgets. Verified 2026-08-30.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click