toolcall() ← all concepts

// concept · agents

Why 20 tool calls cost you 200

No bug, no runaway loop, no expensive model. Just an agent doing exactly what it is supposed to do — and re-reading the entire conversation every single time it does.

// the mechanism

The API is stateless. The model has no memory between calls, so the transcript is the memory — and you send all of it, every step:

The API is stateless — send the full conversation history each time.

Every step of an agent loop re-sends the prior messages, all previous tool results, and the tool definitions. Step one pays for step one. Step twenty pays for steps one through twenty.

The arithmetic, exactly

Say each step adds about 1,000 tokens. Naively you would budget 20 steps × 1,000 = 20,000 input tokens. What you actually send is the running total each time:

1000 × (1 + 2 + … + 20) = 1000 × n(n+1)/2 = 1000 × 210 = 210,000 input tokens ← not 20,000
It is not "the square", and the difference matters. n(n+1)/2 is about HALF of n². At 20 steps that is 210, where "the square" would have you expect 400. It is Θ(n²) as an order of growth, but if someone hears "twenty times twenty" they land on nearly double the real number. Say it climbs far faster than your step count — which is exactly true and needs no hedge.

The same correction applies to the fix. Halving the steps does not quarter the bill: 210 ÷ 55 = 3.82. It only approaches four as n grows. "Cuts it by roughly four" is honest at realistic step counts.

Where the bill is gentler than the token count

The token count grows that way. The bill does not grow quite as steeply, and it would be an overclaim to imply otherwise. Three things soften it:

# prompt caching the stable prefix rebills at ~0.1× # compaction older turns get trimmed or summarised away # fixed overhead system prompt + tool defs do not accumulate — they repeat

So treat the triangular number as the upper bound on tokens, not a quote for your invoice. The shape of the growth is the point; the exact figure depends on your caching and compaction.

The verdict

Cap the loop before you shop for a cheaper model. Step count is the lever with the superlinear effect on what you send — swapping models scales a number that is already growing quadratically, while cutting steps changes the growth itself.

Related: nothing is telling your agent to stop is how you actually cap it, prompt caching is the ~0.1× above, and output tokens covers the other half of the invoice.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click