toolcall() ← all concepts

// concept · workflows

Your retry loop will never clear this 429

Your app gets a 429, waits, tries again. That is the correct response to a rate limit, and it is what the client library does for you without being asked. It is also, for one particular 429, a loop that will fail every single time until the first of next month.

// one status code, three conditions

The errors page lists what a 429 can mean, and only the first item is what most handlers assume:

429 — rate_limit_error: Your organization has hit a rate limit, reached its usage tier's monthly spend cap, or reached a spend limit on the Claude Code workspace. A tier spend-cap 429 has no retry-after header and keeps failing until access resumes.
rate limit RPM / ITPM / OTPM, per model, token bucket retry-after present · clears in seconds tier spend cap Start $500 · Build $1,000 · Scale $200,000 per calendar month no retry-after · clears 00:00 UTC on the 1st, or when the cap is raised Claude Code ws a limit you set on that workspace can return a 429 WITH retry-after — the exception to the exception

Three conditions, one HTTP status, one error.type. The field every retry handler branches on is identical across all of them.

Indistinguishable by type

The rate limits page says so directly:

The error type is rate_limit_error, the same as for a rate limit, but the response has no retry-after header. Retrying, including the SDKs' automatic retries, fails until access resumes.

What "until access resumes" means is spelled out too: "Once you reach your tier's spend cap, API usage pauses until 00:00 UTC on the first day of the next month, unless you request a higher limit sooner." The response body carries the date:

{ "type": "error", "error": { "type": "rate_limit_error", "message": "You have reached your API usage limits: your organization has crossed its monthly API usage threshold, set based on your organization's API tier. You will regain access on 2026-09-01 at 00:00 UTC.", "details": { "error_code": "enforced_spend_limit_reached" } }, "request_id": "req_018EeWyXxfu5pfWkrYcMdjWG" }

So this is not a hidden failure. The message names the day. The problem is that nothing in the shape your code inspects first — status 429, type rate_limit_error — distinguishes a wait of twelve seconds from a wait of three weeks.

The retry nobody wrote

Here is the part that makes it a real outage rather than a curiosity. You may never have written a retry loop at all:

The official SDKs automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default, honoring the retry-after header when present. Each SDK client accepts a maximum-retries option to configure or disable this behavior.

Rate limits are on the list. The spend-cap error is a rate limit as far as the SDK's classifier is concerned, so every request from every user of your app spends three attempts hammering a wall that will not move until the month turns — and then surfaces as a RateLimitError that your own handler, quite reasonably, may retry again. The code doing the damage is code you never saw.

This is not a reason to disable SDK retries. For the real rate limit they are exactly right. It is a reason to catch the exception and look one level deeper before letting anything retry.

Read the field, then branch

The docs give you the discriminator by name:

On the Messages API, error.details.error_code is enforced_spend_limit_reached. Use it to tell this response apart from a rate limit.

Lead with that field. The missing retry-after header is the symptom the docs describe, and a fine second signal, but a future response shape could add a header and the field is the documented contract. In the SDKs, the parsed error body is on the exception:

# Python — the shape is the same in every SDK: catch the typed error, read the body try: message = client.messages.create(...) except anthropic.RateLimitError as e: code = ((e.body or {}).get("error", {}).get("details") or {}).get("error_code") if code == "enforced_spend_limit_reached": # do NOT retry. Alert a human; raise the cap or wait for the date in e.message. raise SpendCapReached(e.message) from e # a real rate limit: the SDK already backed off twice; back off further here

Three things happen at that branch. The request stops being retried. Someone who can raise the cap finds out today rather than at month end. And every other request stops paying for a wall nobody can push through.

The limit you set yourself is a different error

Two neighbours, both verified on the same pages, both worth a branch of their own.

Your own spend limit is a 400, not a 429. "When usage reaches a spend limit you set, requests return HTTP 400 with error type invalid_request_error." The message begins "You have reached your specified API usage limits" (or the workspace variant). Same feeling for your users, different exception class, and the SDK does not auto-retry a 400 — so this one at least fails fast. The exception is the Claude Code workspace, whose limits "can instead receive a 429 that carries a retry-after header".

And the myth this page is not about. A common belief is that a large max_tokens eats into your output-tokens-per-minute budget up front. It does not: "OTPM rate limits are evaluated in real time as output tokens are produced, counting only the actual tokens generated. The max_tokens parameter does not factor into OTPM rate limit calculations, so there is no rate limit downside to setting a higher max_tokens value." If you were lowering max_tokens to avoid 429s, you were solving a problem that does not exist.

Related: output token costs covers what the reply actually costs, which is the other half of a bill that hits a cap; the agent loop covers the loop that multiplies every one of these requests.

Sources: Claude Platform docs — Rate limits (Spend limits → Reaching your spend cap; Rate limits; Cache-aware ITPM); Claude API errors (HTTP errors, SDK retry behaviour). Verified 2026-09-16. Tier caps are Anthropic's own and change; Claude Platform on AWS differs.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click