How to Evaluate AI Agents: Metrics, Test Sets and Human Review

An agent that worked in the demo can still fail one run in three. What to measure, how to build a small test set from your own tasks, which grader to use for what, and a weekly loop a small team can keep up.

7 min read

AI agent evaluation means running an agent on a fixed set of realistic tasks, several times each, and grading both what it produced and how it got there. Measure four things: whether the task succeeded, whether the path was sensible, what it cost in tokens and time, and whether it broke any rule it must never break. Start with twenty to fifty tasks taken from real work, grade with code wherever you can, use a model or a person for the rest, and rerun the set every time you change the model, the prompt or the tools.

This is about evaluating agents: systems that take several steps and call tools. Evaluating an AI feature inside your product, such as a classifier or a summariser, is covered in managing AI projects. Checking a single piece of work an agent hands back is verifying AI-generated work. Evaluation sits between the two: it tells you whether the agent can be trusted with a kind of work at all.

Why agents need their own kind of evaluation

A single model call gives one output to grade. An agent takes many turns, calls tools, changes things, and can reach the right answer by a wasteful or unsafe route. Anthropic’s engineering guide, Demystifying evals for AI agents (opens in a new tab) (January 2026), sets out the vocabulary most teams now use. A task is one test with defined inputs and success criteria. A trial is one attempt at it, and you run several because outputs vary. The transcript, or trajectory, is the full record of a trial: outputs, tool calls, reasoning and intermediate results. The outcome is the state of the world when the trial ends. A grader is the logic that scores some part of it.

What to measure

  • Task success. Did the outcome meet the criteria? Check the state, not the agent’s report: the test suite passes, the record exists with the right values, the file contains what it should.
  • Trajectory quality. Did it choose sensible tools, in a sensible order, without loops, pointless retries or reading things it had no reason to read? A right answer reached by a bad path is a warning about the next task.
  • Cost and latency. Tokens, tool calls and wall-clock time per task. An agent that succeeds at three times the cost of last week’s version has regressed, even if the success rate held.
  • Safety violations. Counted separately and never averaged away: a write outside the allowed area, a secret printed, a send without approval, an instruction followed from content it was only meant to read.

Measure consistency as well as success. The τ-bench paper (Yao et al., 2024 (opens in a new tab)) introduced pass^k, the chance that an agent succeeds on all of k attempts at the same task, and found that even the strongest function-calling agents of the time succeeded on fewer than half of tasks and were far less consistent over eight tries. pass@k asks whether any of k attempts worked, which suits a developer who will pick the best one. pass^k asks whether every attempt worked, which is what matters when the agent runs unattended.

Build a small eval set from real tasks

Do not start with a public benchmark. Start with your own work, because that is what the agent will do. Anthropic’s guide suggests twenty to fifty simple tasks drawn from real failures; early on, differences between versions are large enough that a small set shows them.

  1. Collect tasks the agent has already done or failed at: closed bugs, routine changes, support requests. Each becomes a task with its starting state and a written success criterion.
  2. Add the awkward ones on purpose: an ambiguous request, a task that should be refused, a task whose data contains instructions it must ignore.
  3. Write down what must never happen for each task, as its own check.
  4. Split the set in two. Capability tasks are the hard ones you expect to fail today and hope to pass later. Regression tasks are ones it already passes and must keep passing.
  5. Keep the set in version control beside the agent’s instructions, so a change to either shows up in review.

Write each task the way you would hand it to the agent for real. The same habits that make a task finishable make it testable, as how to write a task for an AI agent explains.

Grading: code, model or person

Each grader type is good at something different. Use the cheapest one that can judge the criterion honestly.

  • Code-based graders are fast, cheap and reproducible: the tests pass, the JSON validates, the record count is right, a forbidden tool was never called. Their weakness is brittleness: they can fail a valid answer that took an unexpected form.
  • Model-based graders handle open-ended output such as a summary, a reply to a customer or a plan. They need a written rubric and a check against human judgement, because language models used as judges have known biases. The MT-Bench study (opens in a new tab) (2023) found strong judges agreed with human preferences about as often as humans agree with each other, and also documented position, length and self-preference biases.
  • Human graders are the reference the others are calibrated against, and the slowest. Use them for a small slice of every run, for anything the rubric cannot pin down, and for checking that the model grader still agrees with people.

Grade what the agent produced rather than the exact path, or you will fail valid solutions that took a different route. Then grade the path separately for the things that matter: safety rules, cost, and obvious waste. OpenAI calls this trace grading (opens in a new tab): assigning scores or labels to the end-to-end record of an agent’s decisions and tool calls. Whatever you use, read a sample of transcripts yourself. Anthropic’s guide is blunt that you will not know whether your graders work unless you read the transcripts and grades from many trials.

Offline and online evaluation

Offline evaluation is the eval set: fixed tasks, run before a change ships, with no users affected. It is what lets you compare two prompts or two models fairly, and it should run on every change to the model, the instructions or the tool list. OpenAI’s agent evals guide (opens in a new tab) suggests the same progression: start from traces while debugging, then move to repeatable datasets and eval runs once you know what success looks like.

Online evaluation is watching the agent at work: success and failure on real tasks, cost per task, refusals, how often a person had to step in, and what reviewers found. It catches what the set does not contain. The two feed each other: every real failure you find online becomes a new offline task, so the same mistake cannot return unnoticed. Collecting the data for the online side is the subject of AI agent observability.

If your agent reads outside content, include adversarial tasks. AgentDojo (opens in a new tab) (2024) pairs 97 realistic tasks with 629 security test cases and measures utility with and without attacks, which is a useful model for a few tasks of your own. The attack itself is described in indirect prompt injection.

A small-team eval loop

  1. Keep the set small and real: thirty tasks, each with a success check, a never rule and an owner.
  2. Run it on every change to the model, the instructions or the tools, three trials per task, and record success, pass^3, cost and violations beside the previous run.
  3. Read five transcripts per run, chosen at random, and note anything the graders missed.
  4. Each week, turn the real failures reviewers found into new tasks, and move capability tasks that now pass reliably into the regression half.
  5. Decide in advance what blocks a change: any safety violation, or a drop in regression success. Write that rule down and let the owner, not the person making the change, relax it.

Keeping the evaluation work visible

Evaluation work goes missing when it lives in someone’s notebook. On a fenbs board it is ordinary work. Each change to an agent’s model, instructions or tools is a task, with the before-and-after numbers in its Testing section: a status of Tested, Partly tested, Failed or Needs owner check, and notes saying what was run and what was not. A real failure becomes a bug linked with relatesTo to the change that caused it. The thresholds that block a change go on the Decisions page with who decided and why. An AI assistant can run the set and report the numbers, but moving a task between lanes is a separate permission, so the move to Completed can stay with a person.

Related

Eval sets for AI product features: managing AI projects. Checking one piece of finished work: verifying AI-generated work. The loop an agent runs, step by step: the steps an AI agent takes.

Questions people ask.

What metrics should I use to evaluate an AI agent?

Task success checked against the final state, the quality of the path it took, cost and latency per task, and safety violations counted separately. Add a consistency measure such as pass^k, the share of tasks it gets right on every one of several attempts.

How many test tasks does an agent eval need?

Twenty to fifty real tasks is enough to start, because early differences between versions are large. Grow the set whenever a real failure turns up, and keep a regression half the agent must always pass.

Can I use an LLM to grade my agent?

Yes, for open-ended output that code cannot check, with a written rubric. Calibrate it against human grades on a sample, and watch for known judge biases such as favouring longer answers, the first option shown, or output from its own model family.

What is the difference between offline and online agent evaluation?

Offline evaluation runs the agent on a fixed test set before a change ships. Online evaluation measures it on real work after release. Offline makes comparisons fair; online finds the failures the test set does not contain, which then become new test tasks.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.