// what an eval is
Two parts: a set of test cases that look like your real traffic (plus the tricky edge cases), and a grader that scores each output. Run it, get a number. Now a prompt tweak is "78% → 84%," not a shrug.
How to grade
Pick the cheapest grader that works:
The workflow
You don't need thousands of cases. Start with a few dozen good ones. Then every bug you find becomes a new case — so it can never come back quietly. Run the cheap graders on every change and track the score over time; a dip is a regression you caught before shipping.
Sources: Hamel Husain — Your AI product needs evals · Anthropic — build evaluations · OpenAI — eval best practices · Zheng et al. — LLM-as-a-judge (MT-Bench)
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
