toolcall() ← all concepts

// concept · evals

One run is a sample, not a result

You changed something, the eval went green, and you shipped. But every case ran exactly once — and the same case, same model, same prompt can come back the other way on the next run.

// it moves even at temperature zero

The first objection is always the same: I set temperature to zero, so it's deterministic. Sampling is one source of variation and turning it off removes that one. It does not make the system reproducible.

The cause is batch invariance, and it is worth getting right because the usual explanation is wrong. It is not mainly floating-point non-associativity: most of an LLM forward pass uses deterministic reduction strategies, and running the same forward pass twice in isolation yields bitwise identical output. The problem is that key kernels are not batch-invariant — the numerics for your request depend on the batch it landed in, and on a shared server that batch is different every time.

# same request, two runs run 1 your prompt + 14 strangers' → one batch run 2 your prompt + 31 strangers' → different numerics greedy decoding diverges from there
This is a property of serving, not of the model. The same work that diagnosed it built batch-invariant kernels achieving bit-identical output across 1,000 runs, at roughly 34% overhead. So it is fixable — it is just not fixed by setting temperature to zero, and almost certainly not fixed on the endpoint you are calling.

Two questions, and most evals ask the first

Once a case can go either way, "did it pass?" stops being one question. Anthropic's agent-eval guidance names both:

pass@k measures the likelihood that an agent gets at least one correct solution in k attempts.
pass^k measures the probability that all k trials succeed. As k increases, pass^k falls, since demanding consistency across more trials is a harder bar to clear.

They answer genuinely different things. One asks can it do this; the other asks will it, every time. A benchmark leaderboard cares about the first. A product lives entirely in the second — your users do not get eight attempts and keep the best one.

# four runs of ONE case ✕ ✓ ✕ ✕ pass@k passes ← at least one worked ✕ ✓ ✕ ✕ pass^k fails ← not every one did

How wide the gap actually gets

This is measured, not argued. τ-bench proposed pass^k precisely to score reliability over repeated trials, and found state-of-the-art agents "quite inconsistent". Two figures from the paper, kept in their own scopes:

# tau-bench (Yao et al., arXiv 2406.12045) single try, overall function-calling agents succeed on <50% retail domain pass^8 < 25% ← passed all eight
Those are different columns. The under-50% figure is the overall single-try rate; the under-25% is retail's pass^8 specifically. Stitching them into one sentence reads better and would be wrong, and the widely-quoted "60% drop" comes from a secondary write-up rather than the paper. The honest headline is retail's own: fewer than a quarter of tasks passed all eight trials.

So run it more than once

The change is small and unglamorous: stop scoring one run per case. Anthropic states the practice plainly — "because model outputs vary between runs, we run multiple trials to produce more consistent results."

# not this score = run(case) # this trials = [run(case) for _ in range(k)] score = all(trials) # and keep the fraction that were all-pass

No one will tell you the right k, and this page will not either — the sources do not give one, and eight is τ-bench's number rather than a rule. What matters is that it is greater than one, that it is the same k every run so scores stay comparable, and that a case which passes sometimes is reported as what it is: unstable, not green.

Not the same as having too few cases

These two get conflated constantly, and they are different axes that compose rather than substitute.

how many cases variance ACROSS cases → sampling error in the score how many runs variance ACROSS trials → is any one case stable?

A thousand cases run once each still tells you nothing about whether any single one of them is repeatable. And running one case a hundred times tells you nothing about coverage. Your eval set is too small to trust covers the first; this page covers the second.

Related: your agent said it worked — check the database is what a single case should assert in the first place, and your eval measures a metric you guessed is where the case should have come from.

Sources: Anthropic — Demystifying evals for AI agents · Yao et al. — τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv 2406.12045) · Thinking Machines Lab — Defeating Nondeterminism in LLM Inference. Verified 2026-09-01.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click