Skip to main content

After building an agent, how do we know it works well, rather than just feeling good?

My AI assistant is meant to respond as me. Previously, when I asked it “Who are you?”, it would answer in that role, introducing itself as me. After DeepSeek updated its LLM, however, the assistant began answering “I am DeepSeek” instead. It was no longer maintaining the identity I expected it to represent. This change made me interested in evaluations: how can we check whether an agent still behaves as intended after an underlying model update?

Agents still need unit tests and end-to-end tests. The additional challenge is that the same task can produce different behavior across runs. A single successful demonstration does not establish reliability, and changes to the model, prompts, tools, or agent harness can introduce regressions.

This journal is based on Anthropic's Demystifying evals for AI agents with a calendar example to connect the concepts.

What an agent evaluation measures

An evaluation gives an agent a task and uses explicit grading rules to assess its behavior and results. For an agent that changes external state, checking its final response alone is not enough: saying “I created the event” does not prove that the event exists.

The main components are:

  • Task: a test case with defined inputs, an initial environment, and success criteria.
  • Trial: one attempt at a task. Repeated trials help estimate how reliably the agent succeeds.
  • Transcript/trace/trajectory: the observable record of a trial, including messages, tool calls, tool results, and any reasoning content or summaries the system exposes. It does not imply access to hidden internal reasoning.
  • Outcome: the final result, such as the calendar state after the trial. For a research agent, the result may instead be a report or another artifact.
  • Grader: logic that evaluates an aspect of the transcript or outcome. A task can have multiple graders.
  • Evaluation suite: a collection of tasks covering particular capabilities or behaviors.
  • Agent harness: the runtime that supplies context, orchestrates model and tool calls, and manages the agent's execution. We evaluate the model and harness together.
  • Evaluation harness: the infrastructure that prepares environments, runs trials, captures evidence, invokes graders, and aggregates results.

A running example: scheduling a bicycle ride

Consider a personal assistant that can check the weather and manage a calendar. We can use one simple request to connect the concepts above.

Task: what should the agent do?

Task #001

User: "If Saturday is not raining, please help me schedule a bicycle ride."

For this example, assume the assistant already knows the user's location, time zone, preferred riding time, and ride duration, and is expected to avoid calendar conflicts and duplicate events. The goal is to check Saturday's forecast and create a suitable calendar event only if no rain is forecast. If rain is forecast, leaving the calendar unchanged and explaining why is also a successful result.

Trial: one attempt at the task

We run the same task multiple times to see whether the agent handles it consistently:

Task #001

Run #1 = Trial 1
Run #2 = Trial 2
Run #3 = Trial 3

Each trial starts with the same weather information and calendar state. Otherwise, an event created in Trial 1 could affect Trial 2, making the results difficult to compare.

One successful trial tells us the agent can handle this example. Repeated trials help us understand how reliably it does so.

Transcript: what happened during the attempt?

For a successful trial with a dry forecast, the execution might look like this:

The transcript records the actual messages, tool calls, and tool results behind these steps. It helps us investigate failures: did the agent query the wrong date, overlook a calendar conflict, or claim success after a tool error?

This is one possible execution path, not a tool sequence every trial must follow exactly.

Outcome: what actually changed?

The final response might say, “Your bicycle ride has been scheduled.” The outcome is whether the correct event actually exists in the calendar.

If rain is forecast, the expected outcome is different: no ride event should be created. Success depends on satisfying the request, not simply on taking an action.

Graders: how do we judge the attempt?

We turn the request into a few concrete checks:

GraderWhat it checks
Weather checkDid the agent obtain the forecast for the correct location and Saturday?
Conditional actionDid it schedule the ride only when no rain was forecast?
Calendar outcomeWhen scheduling was appropriate, was the event created at a suitable time without a conflict or duplicate?
Response accuracyDoes the final response match the tool results and actual calendar state?

These checks can expose a failure that sounds convincing in the final response. For example, if no event exists in the calendar but the agent says the ride is scheduled, both the outcome and response checks should fail.

Suite and harness: how does this become an evaluation?

An evaluation suite collects related tasks: scheduling a ride in dry weather, handling a rainy forecast, or responding when the calendar has no available slot.

The evaluation harness prepares each trial, runs the agent, records the transcript, inspects the calendar, and applies the graders. It then aggregates the results so we can compare behavior across trials or agent versions.

3 Types of Graders

Code-based grader

Like traditional tests, code-based graders can use exact matches, regular expressions, unit tests, static analysis, and checks of the final environment state. For the bicycle ride example, a check can verify whether the expected calendar event actually exists.

These graders are usually fast, inexpensive, and reproducible, so use deterministic checks where possible. However, a check can still reject a valid result if it encodes the wrong assumptions.

Token usage and latency can also be tracked. They are metrics rather than pass/fail checks unless we define a requirement, such as a maximum token budget.

Model-based grader

Some qualities are difficult to judge with code alone. Is the response clear? Is a research report comprehensive? Does it support its claims with evidence?

A common approach is LLM-as-a-Judge: using an LLM to evaluate the output or execution trace against a rubric. For example:

Accuracy: 1–5
Completeness: 1–5
Clarity: 1–5
Groundedness: 1–5

Each dimension needs clear criteria, not just a number. For example, a high groundedness score should mean that the answer's claims are supported by the supplied evidence.

LLM graders handle open-ended answers, but they can be inconsistent or wrong and cost more than simple code checks. Calibrate their judgments against human-reviewed examples, and allow an “Unknown” result when there is insufficient evidence.

Human grader

Human reviewers can assess nuanced results, audit automated grades, and calibrate LLM graders. Their main limitations are cost, time, and difficulty scaling. Clear criteria also matter for human review, since reviewers can disagree.

A practical evaluation may combine:

  • Code graders for verifiable requirements.
  • LLM graders for qualities that require judgment.
  • Human review for calibration and audit.

Choose the combination that fits the task; not every task needs all three.

Capability vs Regression

Capability

Capability evals ask what the agent can do. Include challenging tasks that it cannot yet handle reliably, so there is room to measure improvement. A low initial pass rate can be useful.

Illustrative pass rates:

Easy tasks: 100%
Medium tasks: 80%
Hard tasks: 40%
Extreme tasks: 10%

Regression

Regression evals ask whether the agent still handles tasks it previously handled well. These should aim for a nearly 100% pass rate.

Before a prompt change:

Search: 100%
Calendar: 100%
Weather: 100%

After the change:

Search: 100%
Calendar: 70% <- Investigate a possible regression
Weather: 100%

These numbers are illustrative. Repeated trials help distinguish a persistent decline from ordinary run-to-run variation.

As capability tasks become consistently solvable, they can become part of the regression suite. Add harder tasks to the capability suite so it continues to reveal opportunities for improvement.

pass@k vs pass^k

Running multiple trials lets us ask two different questions:

  • pass@k: what is the probability that at least one of k attempts succeeds?
  • pass^k: what is the probability that all k attempts succeed?
At least one success in three trials:
Trial 1 ❌
Trial 2 ❌
Trial 3 ✅

All three trials succeed:
Trial 1 ✅
Trial 2 ✅
Trial 3 ✅

The second group satisfies both conditions. These examples illustrate the conditions, rather than precisely estimating the underlying probabilities.

Suppose independent attempts at one task each have a fixed 75% success probability:

pass@3 = 1 − (1 − 0.75)^3 ≈ 98.4%
pass^3 = 0.75^3           ≈ 42.2%

More attempts make at least one success more likely, while requiring every attempt to succeed sets a higher reliability bar. These formulas assume independent attempts with the same success probability for that task.

pass@k asks, “Can it succeed within k attempts?” pass^k asks, “Can it succeed consistently across k attempts?”

Build great evals for agents

Start with real tasks

Start with 20–50 real tasks rather than waiting until you have hundreds. Useful sources include:

  1. Real user requests.
  2. Checks already performed during development.
  3. Failures the agent has encountered.

For example, if a calendar agent scheduled an event during an occupied time slot, turn that failure into an eval task. A small suite is a practical starting point; more mature systems may need more tasks to detect smaller improvements.

Each task should have success criteria

Consider this request:

Please schedule a 30-minute meeting for Friday afternoon,
using an available slot in my calendar.

Assuming the date and time zone are clear from context, the checks could be:

✓ A meeting was created.
✓ It falls on Friday afternoon.
✓ It lasts 30 minutes.
✓ It does not conflict with an existing event.

Explicit criteria turn “it feels right” into something we can evaluate. If required information is missing, clarify it rather than grade the agent against a hidden expectation.

Prepare reference solutions

Some failures come from the task or grader rather than the agent. Prepare a known-working solution that passes the graders. This helps confirm that the task is solvable and the checks are configured correctly.

A reference solution is an example of a valid result, not necessarily the only acceptable way to solve the task.

Build balanced datasets

Test both when the agent should search and when it should not. For example:

Should search:
- Today's weather
- Latest news
- Current stock price

Usually does not need search:
- Who founded Apple?
- What is 1 + 1?
- What is a Python list?

If we only test whether the agent searches when it should, we may encourage it to search for almost everything. Test both whether it searches when needed and whether it avoids unnecessary searches.

Build isolated environments

Each trial should start from a clean environment. Leftover files, database changes, cached information, or Git history from previous trials can affect the result.

For the bicycle ride example, reset the calendar before each trial so an event created in an earlier attempt does not influence the next one. The evaluation environment should also reflect how the agent operates in production closely enough to make the results useful.

Prioritize checking the Outcome

Check what the agent accomplished rather than requiring one exact sequence of tool calls. A different approach may still produce a valid result.

For example, verify that the correct calendar event exists instead of merely checking whether the agent called the creation tool. Process checks still matter when the process itself is a requirement, such as verifying identity before a transfer.

Check the transcripts

A score such as “60% success” does not explain what to improve. Read transcripts alongside grades to understand what happened and whether the judgment was fair.

Tool use: Was the wrong tool selected or called incorrectly?
Context: Was necessary information missing?
Instructions: Did the agent misunderstand or ignore a requirement?
Environment: Did a tool or environment failure prevent completion?
Grader: Was a valid result rejected by an unreasonable check?

Inspect successful trials too: an apparent pass may reveal a weak grader. Use the evidence to decide whether to improve the agent, the task, or the evaluation itself.

The roadmap is:

A good first agent eval does not need to be large, but it must be realistic, clear, and verifiable.

Fit with other methods

Automated evals measure performance on a predefined set of tasks. They cannot cover every real user situation, so pair them with other methods.

Production monitoring

Monitoring reveals how real users use the agent and what failures occur in production. Turn relevant new failures into eval tasks so the suite continues to reflect actual usage.

A/B testing

A/B testing compares versions with real users. A version that scores better in offline evals may not improve the user experience in practice.

It requires real traffic and enough time to produce meaningful results, which can take days or weeks depending on the experiment.

User feedback

Many agent products let users report problems with a response or action. This feedback can reveal issues that automated checks miss and provide useful cases for the eval suite.

No single method catches every problem:

MethodKey question
Automated EvalsHow well does the agent perform on our predefined tasks?
Production MonitoringWhat problems does the agent actually encounter in production?
A/B TestingIs the new version better for real users?
User FeedbackWhat problems do users actively report?
Transcript ReviewWhat happened during a success or failure?
Human EvaluationHow do human experts assess the quality?

These methods feed into an ongoing improvement cycle. The diagram is a simplified loop; not every change requires an A/B test.

For my own assistant, a simple regression task would be to check whether it still responds as me after a model update. That would turn the behavior I care about into something I can test repeatedly.

Source

Demystifying evals for AI agents — Anthropic