// the mechanism
The API is stateless. The model has no memory between calls, so the transcript is the memory — and you send all of it, every step:
Every step of an agent loop re-sends the prior messages, all previous tool results, and the tool definitions. Step one pays for step one. Step twenty pays for steps one through twenty.
The arithmetic, exactly
Say each step adds about 1,000 tokens. Naively you would budget 20 steps × 1,000 = 20,000 input tokens. What you actually send is the running total each time:
n(n+1)/2 is about HALF of n². At 20 steps that is 210, where "the square" would have you expect 400. It is Θ(n²) as an order of growth, but if someone hears "twenty times twenty" they land on nearly double the real number. Say it climbs far faster than your step count — which is exactly true and needs no hedge.The same correction applies to the fix. Halving the steps does not quarter the bill: 210 ÷ 55 = 3.82. It only approaches four as n grows. "Cuts it by roughly four" is honest at realistic step counts.
Where the bill is gentler than the token count
The token count grows that way. The bill does not grow quite as steeply, and it would be an overclaim to imply otherwise. Three things soften it:
So treat the triangular number as the upper bound on tokens, not a quote for your invoice. The shape of the growth is the point; the exact figure depends on your caching and compaction.
The verdict
Cap the loop before you shop for a cheaper model. Step count is the lever with the superlinear effect on what you send — swapping models scales a number that is already growing quadratically, while cutting steps changes the growth itself.
Related: nothing is telling your agent to stop is how you actually cap it, prompt caching is the ~0.1× above, and output tokens covers the other half of the invoice.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
