toolcall() ← all concepts

// concept · agents

Nothing is telling your agent to stop

There is no "finished" signal coming from your tools. The loop ends because the model says it is done — and if it never says so, only your own code is left to stop it.

// who ends the loop

Every model response carries a stop reason, and it says one of two useful things: I want to call a tool, or I'm done. Your harness reads that field and either runs the requested tools and goes round again, or breaks.

# the agentic loop, in full while stop_reason == "tool_use": run the tools append the results as tool_result blocks ask again # stop_reason == "end_turn" → the model is done

A tool cannot end the loop. A tool returns a result, that result is appended to the context, and the model decides what it means. There is no "terminate" value in the protocol — nothing a tool can return that means stop now.

The field names here are Anthropic's. stop_reason, end_turn and tool_use are Claude API vocabulary; other providers express the same idea with different names. The mechanism is general — a model that signals its own completion — but do not go looking for these exact fields in someone else's SDK.

So the guard is your code's job

If the model never says it is done, nothing else will. The backstop is a maximum-iteration counter around your own loop — a harness concern, not a model or tool concern:

for step in range(MAX_STEPS): ... # and decide what happens when you hit it

Be clear about the size of the risk, though. Models usually do terminate; the cap exists for the tail case — a tool that errors in a way that keeps inviting a retry, or a goal the model cannot satisfy and will not abandon. It is a seatbelt, not a daily occurrence.

And a cap is not the only backstop worth having. Timeouts, spend limits, and a human-approval gate on anything destructive all cover failures an iteration counter does not.

A cap the model can't see cuts it off mid-work

Here is the part that bites. A hard ceiling like max_tokens is enforced but never surfaced to the model. It is working, it has no idea a limit is approaching, and then it simply stops — mid-sentence, mid-edit, mid-plan. You do not get a wrapped-up answer. You get a truncated one.

# invisible ceiling max_tokens enforced, model never told → cut off # visible allowance a budget it can see model paces itself → finishes

The durable principle is the one worth keeping: a limit the model can see changes how it behaves; a limit it cannot see only changes where it gets severed. Given a budget it can observe, a model spends it — prioritising, then summarising what it got done.

Anthropic ships this as task budgets — a token allowance for the whole agentic loop, with a countdown the model observes while it works, explicitly contrasted with max_tokens in the API docs. It was in beta with a 20,000-token minimum when this was written, so check the current state before you build on it.

The rule

Cap the loop in your code, and give the model a limit it can see. The first stops a runaway; the second is the difference between an agent that finishes and one that gets cut off.

Related: an AI agent is just a loop is the shape this sits inside, and why a long run costs so much more than you'd guess is the other reason to cap the step count.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click