// it already does this
Parallel tool use is not a feature you turn on. Anthropic's tool-use documentation states it twice, and both times as the default:
tool_use blocks.So a single assistant turn can arrive carrying three requests, and it will, without you configuring anything. What you control is the shape of the reply:
The instruction has two halves
Here is the sentence the whole topic rests on, and it is worth reading as two separate obligations rather than one:
tool_result blocks in a single user message — splitting them across multiple messages silently trains Claude to stop making parallel calls.Half one is your code's job, and nothing does it for you. Receiving three tool_use blocks in one turn does not run anything in parallel. If your handler is a for loop over the batch, it takes exactly as long as answering them one at a time would have. That is the half people assume is free, and it isn't.
Half two holds however your code is written, which is why it is the one worth building the habit around.
What batching saves unconditionally
The API is stateless. The model has no memory between calls, so the only way to obtain the next action is another request carrying the entire conversation. Count those requests and the difference is structural rather than a matter of how fast your tools are:
For n tools it is n+1 turns serially and always 2 when batched. And each of those turns re-reads the whole conversation before producing anything, which is usually slower than the tool it was waiting on — so the extra turns are rarely the cheap part.
The half that nobody sees coming
Answer a batch one message at a time and the model stops asking for batches. That is what "silently trains" means in the sentence above — and the scope of it matters, because the word suggests something much bigger than what happens.
Nothing is retrained and nothing persists. The model is conditioning on this conversation's history: it sees its parallel request answered serially and adapts. Your account is unaffected, other conversations are unaffected, and a fresh conversation starts batching again.
Which is exactly why it is hard to catch. There is no error, no warning, and no config to inspect — the run simply gets slower the longer it goes, because the ask shrinks from three to two to one.
Two rules, and one caveat
Run them at the same time, and send every result back in one message. Both halves, in that order.
A failed tool still counts as a result. Dropping it leaves the batch incomplete, which makes the turn malformed:
And the caveat: parallel is not always right. Tools that write, or that depend on each other's output, must not run concurrently — that is what disable_parallel_tool_use exists for. If your tools write, making them idempotent is the prerequisite, not the batching. Related: the model is stateless between calls is the fact the turn count rests on, and what the reply actually bills covers the other half of what a turn costs you.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
