toolcall() ← all concepts

// concept · evals

Your eval measures a metric you guessed

Where did the things your eval scores actually come from? Almost always: someone sat down and made a list. Helpfulness, tone, accuracy, relevance. Meanwhile the thing that actually breaks is that it picked the wrong customer record — which none of those score at all.

// generic metrics are worse than useless

The failure is not that a guessed metric is unhelpful. It is that it moves. You tune a prompt, helpfulness goes from 3.8 to 4.1, and you have learned nothing about the defect your users are actually hitting — but you have a number that says you improved.

The practitioner literature is blunt about it: error analysis exists so that the metrics you develop are grounded in real application behaviors instead of counter-productive generic metrics. Counter-productive, not merely weak.

invented helpfulness · tone · accuracy · relevance observed wrong customer record · answered in the wrong language · invented an order number

One of those lists could belong to any product. The other could only belong to yours, which is the entire point.

Read the runs before you write the metric

Start with real traffic, not a spec. Open actual runs and, for each one, write down what went wrong in ordinary words — no categories yet, no scoring, no schema. The discipline is to describe rather than classify, because classifying early means you only ever find the categories you brought with you.

# one note per run, plain English run 041 "looked up the wrong customer, then answered confidently about their order" run 042 "gave up after one failed tool call"
How many? The rule of thumb is to review at least 100, with a stopping heuristic: if roughly 20 traces in a row turn up no new category, you can stop. It is a rule of thumb from practice, not a finding — but the shape of it matters more than the number, because it tells you when you are done rather than leaving it to stamina.

Group them, then count

Only now do categories appear, and they come out of the notes rather than into them. Cluster the free-form notes into a taxonomy — typically five to ten failure modes — then count how often each one occurs and rank engineering effort by frequency.

# your failure modes, counted wrong record ████████████████ 31 wrong language █████████ 18 made-up data ██████ 12 gave up early ████ 9 ignored policy ███ 6

The two passes have names in the research literature — open coding for the free-form notes, axial coding for the grouping — but nothing about the method requires the vocabulary. Write down what broke; put like with like; count.

Then build the eval for the top of that list. That metric came from your product instead of your imagination, and when it moves you know exactly what moved.

A model can help — after you have done thirty

The obvious objection is cost, and it has a real answer. An LLM can assist with the clustering once 30–50 traces have been coded by hand. What it cannot do is the initial pass, because open coding depends on domain knowledge — knowing that this particular reply is subtly wrong for this particular customer is the whole skill.

human read + write what went wrong ← needs your domain model group the notes into modes ← after the first 30-50 human decide what to fix first

So the expensive part is bounded. You are not signing up to hand-label a hundred runs forever; you are signing up to genuinely read the first few dozen.

It also catches your eval being wrong

An underrated second payoff. Reading transcripts does not only tell you how the system failed — it tells you when it didn't:

When a task fails, the transcript tells you whether the agent made a genuine mistake or whether your graders rejected a valid solution.

You cannot get that from a score, by construction. A number cannot tell you it is measuring the wrong thing, which is why "you won't know if your graders are working well unless you read the transcripts and grades from many trials" is stated as a requirement rather than a suggestion.

Related: an eval turns "seems fine" into a number ends on adding every prod bug to your set — this is the deliberate version of that, done a hundred at a time before you decide what to measure. Check the database, not the reply is the wrong-object mistake where this is the wrong-property one, and one run is a sample covers how many times to run whatever you land on.

Sources: Husain & Shankar — AI Evals for Engineers & PMs · Anthropic — Demystifying evals for AI agents. Verified 2026-09-01.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click