// it moves even at temperature zero
The first objection is always the same: I set temperature to zero, so it's deterministic. Sampling is one source of variation and turning it off removes that one. It does not make the system reproducible.
The cause is batch invariance, and it is worth getting right because the usual explanation is wrong. It is not mainly floating-point non-associativity: most of an LLM forward pass uses deterministic reduction strategies, and running the same forward pass twice in isolation yields bitwise identical output. The problem is that key kernels are not batch-invariant — the numerics for your request depend on the batch it landed in, and on a shared server that batch is different every time.
Two questions, and most evals ask the first
Once a case can go either way, "did it pass?" stops being one question. Anthropic's agent-eval guidance names both:
pass^k measures the probability that all k trials succeed. As k increases, pass^k falls, since demanding consistency across more trials is a harder bar to clear.
They answer genuinely different things. One asks can it do this; the other asks will it, every time. A benchmark leaderboard cares about the first. A product lives entirely in the second — your users do not get eight attempts and keep the best one.
How wide the gap actually gets
This is measured, not argued. τ-bench proposed pass^k precisely to score reliability over repeated trials, and found state-of-the-art agents "quite inconsistent". Two figures from the paper, kept in their own scopes:
So run it more than once
The change is small and unglamorous: stop scoring one run per case. Anthropic states the practice plainly — "because model outputs vary between runs, we run multiple trials to produce more consistent results."
No one will tell you the right k, and this page will not either — the sources do not give one, and eight is τ-bench's number rather than a rule. What matters is that it is greater than one, that it is the same k every run so scores stay comparable, and that a case which passes sometimes is reported as what it is: unstable, not green.
Not the same as having too few cases
These two get conflated constantly, and they are different axes that compose rather than substitute.
A thousand cases run once each still tells you nothing about whether any single one of them is repeatable. And running one case a hundred times tells you nothing about coverage. Your eval set is too small to trust covers the first; this page covers the second.
Related: your agent said it worked — check the database is what a single case should assert in the first place, and your eval measures a metric you guessed is where the case should have come from.
Sources: Anthropic — Demystifying evals for AI agents · Yao et al. — τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv 2406.12045) · Thinking Machines Lab — Defeating Nondeterminism in LLM Inference. Verified 2026-09-01.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
