// two different facts
Agent evaluation has two objects, and they are easy to collapse into one. Anthropic's engineering write-up names them separately:
An eval that reads the transcript is grading what the agent said it did. An eval that reads the outcome is grading what happened. Those diverge more often than anyone is comfortable with, because a language model is extremely good at producing the sentence that should follow a success — whether or not the success occurred.
The flight that was never booked
The canonical example, near-verbatim from the source, is worth keeping in your head as an image:
Same for support: an agent can claim success verbally while never processing the refund. Grading on conversational quality misses genuine task completion — it scores the story, not the fact.
Make the test go and look
The fix is not a better rubric. It is a different assertion. Your test case ends by querying the environment it was supposed to change.
The rule that follows: a missing row is a failure no matter how confident the reply was. That sounds obvious written down, and it is the single most common thing missing from agent evals in the wild, because asserting on text is so much easier to write than standing up an environment you can inspect afterwards.
Don't score pass or fail
The second half, and the part people skip. A binary score throws away the information you most need while debugging.
Both of those runs are red under pass/fail, and one of them is nearly working. Partial credit is what lets you see a regression move.
What this doesn't say
It does not say transcripts are useless. The transcript is how you find out why it failed — which tool errored, what it did instead, where the reasoning went sideways. It is just not the proof that it succeeded. Keep it beside the row, not instead of it. Assessing the trajectory has its own value too: it can reveal unsafe reasoning even on runs where the outcome came out fine.
Related: an eval turns "seems fine" into a number is the prerequisite — this page assumes you already have one. Your eval set is too small to trust covers how many cases you need; this covers what a single case should assert. And grading with a model is for the genuinely subjective calls — note that the check on this page isn't one. A row exists or it does not.
Source: Anthropic — Demystifying evals for AI agents. Verified 2026-08-30.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
