All posts

AI agent evaluation

/AI/4 min read

Judging an agent by whether it finished is not enough, because per-step reliability compounds. A 95%-per-step agent completes a twenty-step task about a third of the time.

An agent does not produce one answer. It produces a sequence of decisions — which tool, which arguments, what to do with the result — and then an answer at the end.

Scoring only the answer therefore discards most of what happened, including whether it was earned or lucky.

Why the steps have to be measured

Because per-step reliability compounds, and the compounding is brutal.

per step5 steps10 steps20 steps50 steps
0.99999.5%99.0%98.0%95.1%
0.9995.1%90.4%81.8%60.5%
0.9577.4%59.9%35.8%7.7%
0.9059.0%34.9%12.2%0.5%
0.8032.8%10.7%1.2%0.0%
RELIABILITY COMPOUNDS 0.999 0.99 0.95 0.90 0 20 40 steps in the task 100% 0
The difference between 0.95 and 0.99 per step barely shows at five steps and dominates at fifty.

A 95% success rate per step sounds respectable. Over twenty steps it finishes 35.8% of the time. To complete twenty steps nine times out of ten you need to be right on 99.5% of individual steps.

This is why a final-outcome number is not enough on its own. Two agents can both complete 60% of tasks while having quite different per-step reliabilities, and the one with the better steps will pull ahead as tasks get longer — which is exactly what happens as you deploy it on real work.

What to measure

The outcome. Did it finish the task correctly? Simple to define, matches what a user cares about, and says nothing about why a failure happened.

The trajectory. Did it take sensible steps in a sensible order? This is where debugging lives, and where luck gets separated from skill. An agent that reached the right answer without calling the tool that would have told it the answer has guessed, and it will not guess right next time.

The trouble with trajectories is that there is rarely one correct path. Rather than demanding an exact sequence, it is usually more useful to check whether the required steps all happened (in order or not), what fraction of the steps taken were useful, and what fraction of the necessary ones were missed.

The tool calls specifically. Was the right tool picked, were the arguments valid, did the call succeed, was the result used correctly, and was a non-existent tool ever invented. These are structured and machine-checkable, and they are where a large share of failures actually are.

One run tells you nothing

Agents are not deterministic. The same task twice produces different trajectories, and sometimes different outcomes.

runsinterval on the success rate
1±78.4 pts
3±45.3 pts
10±24.8 pts
30±14.3 pts
100±7.8 pts
250±5.0 pts

Three runs pins the success rate to about ±45 points, which is to say not at all. Telling an 80% agent from a 70% one needs roughly 145 runs of each.

Which puts a hard floor under the cost of evaluating a change. If a prompt tweak is reported to have improved success from 70% to 80% on five runs, that is noise.

The practical compromise is a fixed test set run several times, reporting the average and the worst case. For anything that takes real actions, the worst case is the number that matters.

What else has to be tracked

Success rate alone will happily approve an agent that is unusable.

Steps per task. Fewer is better, and a rising step count is often the first sign of a prompt change gone wrong.

Cost and latency per task. An agent that succeeds 95% of the time at four dollars and ninety seconds a task may be worse than one at 88%, twenty cents and twelve seconds. Nothing in a quality metric surfaces that.

Recovery. When a tool errors, does the agent adapt or collapse? Real tools fail, so this is a property worth measuring on purpose by injecting failures.

Loops. The same call repeated forever. Cap steps and count how often the cap is hit — that number should be near zero and is easy to forget to look at.

Testing safely

An agent under evaluation is taking real actions. Send an email during a test run and the email is sent.

So evaluation runs against a sandbox: fake tools with the same interfaces, seeded data, a database that gets reset. This also fixes reproducibility — an agent evaluated against the live web is being scored partly on what the web did that day.

And build the test set from real tasks, including the ones that go wrong: ambiguous requests, tools that fail, inputs designed to push it somewhere it should not go. Easy tasks confirm the agent works on easy tasks.

The short version

  • An agent produces a sequence of decisions, so scoring only the final answer discards most of the evidence.
  • Per-step reliability compounds: 95% per step is 35.8% over twenty steps.
  • Finishing twenty steps nine times in ten needs 99.5% per step.
  • Outcome says whether it worked; trajectory says why, and separates skill from luck.
  • Tool selection, arguments and result handling are structured and cheap to check automatically.
  • Agents are non-deterministic — three runs measures nothing; distinguishing 80% from 70% takes about 145.
  • Track steps, cost, latency, recovery from tool failures, and how often the step cap is hit.
  • Evaluate in a sandbox, because a test run takes real actions and the live world is not reproducible.