toolcall() ← all concepts

// concept · workflows

Sonnet 5: a third off, or an eighth?

Sonnet 5 costs a third less per token than Sonnet 4.6. That is true, and it is on the price list. But the token itself changed size between the two models, so the discount on the price list is not the discount on your text. Here is the arithmetic, what Anthropic measured, and the check that costs nothing.

// the unit changed

Claude uses a newer tokenizer from Opus 4.7 on, and it cuts the same text into more pieces. The token counting page puts it plainly:

Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer. The same input text produces approximately 30 percent more tokens than on earlier models. The exact increase depends on the content and workload shape.

Sonnet 5 is on the new one; Sonnet 4.6 is on the old one ("Claude Sonnet 4.6 and earlier models use the previous tokenizer"). So a Sonnet 4.6 to Sonnet 5 move crosses the boundary. The Opus 5.5 migration guide gives the range: the tokenizer "may use roughly 1x to 1.35x as many tokens when processing text compared to models before Claude Opus 4.7 (up to ~35% more, varying by content)". About 30% is typical, not a constant.

Think of it as a pizza. The price per slice fell, but the new cutter slices the same pizza into more pieces.

A third off the list, about an eighth off your text

Anthropic's own Sonnet 5 page says the cut does not carry straight through:

Claude Sonnet 5 is priced at $2 per million input tokens and $10 per million output tokens, lower per-token pricing than Claude Sonnet 4.6's $3/$15. Because the new tokenizer produces approximately 30% more tokens for the same text, the cost of an equivalent request does not drop in direct proportion to the per-token prices when comparing with Claude Sonnet 4.6.

The $2/$10 price is permanent: it "is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur." Our arithmetic for the same text, at list price:

# our math, list prices, same text input Sonnet 4.6 $3.00 / M Sonnet 5 $2.00 × 1.30 = $2.60 ≈ 13% less output Sonnet 4.6 $15.00 / M Sonnet 5 $10.00 × 1.30 = $13.00 ≈ 13% less # at the migration guide's upper figure 1.35× $2.00 × 1.35 = $2.70 vs $3.00 10% less

The multiplier applies to text you send and text you get back alike, so the ratio is the same on both sides: roughly an eighth off, not a third. Your own ratio decides the exact number.

Still cheaper, per solved task

None of this means the new model costs you more. Anthropic's cost guide makes the opposite point, and names the unit to compare on:

Compare on cost per solved task, not per token: the same text costs about 30% more tokens on Claude Opus 4.7 and later, so a per-token comparison makes the newer models look more expensive by construction.

For this exact move it measured: "Sonnet 5's saving comes from its lower per-token price, which more than offsets the extra tokens it uses per task compared with Sonnet 4.6: 15% less per solved task for 5 more points." That is one benchmark, at shipped defaults and list rates, but it is Anthropic's own and it points the right way.

And the general pattern, in its words: "in Anthropic's measurements each newer model solved at least as many tasks as the one before it, usually for less per solved task". Usually is doing work there. The same page shows Opus 4.8 to Opus 5 solving "12 more points of tasks at 21% more per solved task", and a Fable 5 to 5.1 upgrade costing 41% more per task on DeepResearch Bench II at high effort. Measure your own workload.

The check costs nothing

Three lines, from the token counting docs:

To measure the difference for your workload, count the same request twice, once with your current model and once with the model you plan to move to, and compare the two input_tokens values.
# Python: same request, two models for model in ("claude-sonnet-4-6", "claude-sonnet-5"): n = client.messages.count_tokens(model=model, messages=messages, system=system, tools=tools) print(model, n.input_tokens) # your ratio = new / old. Then compare cost per SOLVED task on your evals, not per token.

"Token counting is free to use but subject to requests per minute rate limits" based on your usage tier. Swap in the model IDs you actually call.

Where the counter stops. It counts the prompt, not the reply. On Sonnet 5 adaptive thinking is on by default, and max_tokens "is a hard limit on total output (thinking plus response text)", so the output side of the bill is not in that number. The count "is an estimate" and "might differ by a small amount". And it refuses a few inputs the Messages API accepts: server tools such as web search, web fetch and code execution (the advisor tool is the exception), the MCP connector, and image or document blocks with a url or file source (send those as base64 to count them).

When the cut is the cut

The gap only opens when a move crosses the Opus 4.7 boundary. Opus 5 to Opus 5.5 does not: both are on the newer tokenizer, which "later Opus models, including Claude Opus 5.5, also use". Opus 5.5 "costs $4 USD per million input tokens and $20 USD per million output tokens, below Claude Opus 5's $5 and $25", so that 20% per-token cut is a 20% cut on the same text.

One knock-on worth knowing: a million tokens now holds fewer words. "1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens." The pricing page's rule of thumb of 4 characters or 0.75 words per token describes the older tokenizer.

Related: output token costs covers why the reply side of the bill is the bigger one; agent run cost covers how a loop multiplies every request; context window limits covers what fits in one call.

Sources: Claude Platform docs: Token counting; Pricing; What's new in Claude Sonnet 5; Claude Sonnet 5 overview; Optimizing for cost and intelligence; Migrating to Claude Opus 5.5; What's new in Claude Opus 5.5. Verified 2026-09-26. The 13% and 10% figures are our arithmetic at list prices; Anthropic's is the 15% per solved task.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click