// generic metrics are worse than useless
The failure is not that a guessed metric is unhelpful. It is that it moves. You tune a prompt, helpfulness goes from 3.8 to 4.1, and you have learned nothing about the defect your users are actually hitting — but you have a number that says you improved.
The practitioner literature is blunt about it: error analysis exists so that the metrics you develop are grounded in real application behaviors instead of counter-productive generic metrics. Counter-productive, not merely weak.
One of those lists could belong to any product. The other could only belong to yours, which is the entire point.
Read the runs before you write the metric
Start with real traffic, not a spec. Open actual runs and, for each one, write down what went wrong in ordinary words — no categories yet, no scoring, no schema. The discipline is to describe rather than classify, because classifying early means you only ever find the categories you brought with you.
Group them, then count
Only now do categories appear, and they come out of the notes rather than into them. Cluster the free-form notes into a taxonomy — typically five to ten failure modes — then count how often each one occurs and rank engineering effort by frequency.
The two passes have names in the research literature — open coding for the free-form notes, axial coding for the grouping — but nothing about the method requires the vocabulary. Write down what broke; put like with like; count.
Then build the eval for the top of that list. That metric came from your product instead of your imagination, and when it moves you know exactly what moved.
A model can help — after you have done thirty
The obvious objection is cost, and it has a real answer. An LLM can assist with the clustering once 30–50 traces have been coded by hand. What it cannot do is the initial pass, because open coding depends on domain knowledge — knowing that this particular reply is subtly wrong for this particular customer is the whole skill.
So the expensive part is bounded. You are not signing up to hand-label a hundred runs forever; you are signing up to genuinely read the first few dozen.
It also catches your eval being wrong
An underrated second payoff. Reading transcripts does not only tell you how the system failed — it tells you when it didn't:
You cannot get that from a score, by construction. A number cannot tell you it is measuring the wrong thing, which is why "you won't know if your graders are working well unless you read the transcripts and grades from many trials" is stated as a requirement rather than a suggestion.
Related: an eval turns "seems fine" into a number ends on adding every prod bug to your set — this is the deliberate version of that, done a hundred at a time before you decide what to measure. Check the database, not the reply is the wrong-object mistake where this is the wrong-property one, and one run is a sample covers how many times to run whatever you land on.
Sources: Husain & Shankar — AI Evals for Engineers & PMs · Anthropic — Demystifying evals for AI agents. Verified 2026-09-01.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
