AI Agent Evaluation Harness: Trajectory Evals in Production

AI Agent Evaluation Harness: Trajectory Evals in Production

AI Agent Evaluation Harness: Trajectory Evals in Production

A support agent passes your demo, passes your twenty hand-picked test conversations, and ships. Two weeks later a prompt tweak that “only clarified tone” quietly breaks refund handling for one in five customers, and nobody notices until the complaints arrive. The agent never crashed. Every individual response looked fluent. The failure lived in the sequence of decisions, tool calls, and state changes that no one was measuring.

This is the central difficulty of AI agent evaluation. A chat model produces one answer you can grade. An agent produces a trajectory: a chain of reasoning steps, tool invocations, and environment mutations whose quality can only be judged by what ended up true in the world and by how the agent got there. Teams that treat agents like chatbots end up with eval suites that are green while production burns.

By the end of this article you will know how to separate outcome grading from trajectory grading, how to build a simulated environment so trials are repeatable, where an LLM judge helps and where it lies, how to compute reliability with pass^k instead of a single pass rate, and how to wire all of it into a continuous integration (CI) gate with a reference architecture.

What this covers: the evaluation vocabulary, a reference harness architecture, outcome versus trajectory graders, simulated users and sandboxes, LLM judge calibration, the statistics of small suites, CI regression gating, failure modes of the harness itself, and a practical rollout checklist.

Context and Background

Agent benchmarks matured quickly between 2023 and 2026. SWE-bench, introduced by Jimenez et al., assembled 2,294 software engineering problems from real GitHub issues across 12 Python repositories, and at publication the best model, Claude 2, resolved only 1.96% of them (arXiv 2310.06770). That benchmark established the template that most agent evals still follow: give the agent an environment, let it act, and grade the final state with executable tests rather than by reading its prose.

The second landmark was tau-bench from Sierra, which moved the target from code to conversation. It simulates dynamic conversations between a language-model user and an agent equipped with domain tools and policy guidelines, and it introduced pass^k to measure consistency over repeated trials. The authors reported that even a state-of-the-art function-calling agent such as GPT-4o succeeded on fewer than 50% of tasks and had pass^8 below 25% in the retail domain (arXiv 2406.12045). Its successor, tau2-bench, added a dual-control setting in which the user also acts on a shared environment (arXiv 2506.07982). For a comparison of the public benchmarks themselves, see our overview of SWE-bench, GAIA, and tau-bench.

Public benchmarks answer a different question than the one you face. They tell you how a model ranks in general. They cannot tell you whether your agent, with your system prompt, your tool schemas, and your policy documents, still behaves correctly after Tuesday’s deployment. That requires a private harness built from your own failures.

Anthropic’s engineering guidance on agent evals gives a useful shared vocabulary, and this article adopts it (Demystifying evals for AI agents). A task is a single test with defined inputs and success criteria. A trial is one attempt at a task; because model outputs vary, you run several. A grader is logic that scores some aspect of performance. A transcript or trace is the full record of outputs, tool calls, and reasoning. An outcome is the final state of the environment, such as whether a reservation row exists in a database. The harness is the infrastructure that runs agents against tasks, records everything, and aggregates scores.

Two further terms matter. Capability evals ask what the agent can do and start at low pass rates, giving you headroom to climb. Regression evals ask whether things that used to work still work, and they should sit near 100%. Most teams need both, and they need to keep them in separate suites, because a regression suite polluted with hard aspirational tasks teaches everyone to ignore red builds.

There is a related discipline for retrieval pipelines, covered in our piece on RAG evaluation metrics and faithfulness. Agents inherit those problems and add action, state, and multi-turn interaction on top. If you already log traces for observability, the LLM observability and LLMOps architecture post describes the instrumentation layer that the harness described here consumes.

Reference Architecture for an Agent Evaluation Harness

An agent evaluation harness is a pipeline that replays versioned tasks against the agent in a fresh sandbox, records the full trace and the final environment state, scores both with a mix of code, rule, and model graders, and publishes pass^k and confidence intervals to a CI gate. Its defining property is that every trial is isolated, repeatable, and gradeable without a human in the loop.

AI agent evaluation harness reference architecture with sandbox, simulated user, trace store, graders, and CI gate

Figure 1: Reference architecture for an AI agent evaluation harness, from versioned task suite to CI gate.

Figure 1 shows the components. A versioned task suite feeds a runner. For each trial the runner provisions a fresh sandboxed environment seeded from a known state, starts a simulated user if the task is conversational, and launches the agent under test. The agent’s messages and tool calls stream into a trace store. When the run ends, a snapshot of the environment is captured. Graders consume both the trace and the snapshot, and results land in a score ledger that powers dashboards and the CI gate.

The task suite is code, not a spreadsheet

Each task is a small structured record: an identifier, the initial environment seed, the user goal or scripted inputs, the policies the agent must obey, and the grading specification. Store these in git next to the agent’s prompts and tool schemas so one commit captures a coherent version of “the agent plus its tests”. A task that cannot be re-run deterministically from its record is a liability, because you will not be able to tell whether a regression came from the agent or from drift in the test.

Source tasks from real failures. Anthropic’s guidance recommends starting with 20 to 50 realistic tasks drawn from actual failures instead of waiting to assemble hundreds. That advice matches the economics: early in an agent’s life, each fix has a large effect size, so a small suite already discriminates between good and bad changes. Grow the suite by turning every production incident into a new task, a loop we revisit in the CI section.

Isolation per trial

Each trial must begin from a clean environment. Shared state, whether cached data, leftover files, or a database that earlier trials mutated, introduces correlated failures that have nothing to do with agent capability. If trial 3 fails because trial 2 left a lock file behind, your pass^k is measuring your infrastructure. For code-executing agents this means containers or microVMs, and the trade-offs between them are laid out in our comparison of Firecracker, gVisor, and Kata for agent sandboxes. For tool-API agents it usually means an in-memory or ephemeral database reset from a seed file, with mocked or recorded versions of external services.

Isolation has a second purpose: it makes parallelism safe. A harness that can run 50 tasks times 8 trials concurrently finishes a regression suite in the time of its slowest trial rather than the sum, and that difference decides whether engineers run evals before merging or merge first and hope.

Trace capture is the foundation

Every layer above the runner depends on a faithful trace. Record each model request and response, each tool call with arguments and returned payload, token counts, latencies, and any error. Use a stable schema, and ideally an open one such as OpenTelemetry spans carrying GenAI semantic attributes, so the same trace format serves offline evals and production monitoring. When the evaluation trace and the production trace share a schema, you can run the same graders on both, which is the cheapest path to production-grounded evals.

The score ledger

Do not overwrite results. Append each trial’s scores with the commit hash, model identifier, prompt version, suite version, and grader version. When a score moves, you need to answer which of those five things changed. A ledger also lets you compute the baseline for any past commit and diff against it, which the CI gate depends on.

Outcome Grading Versus Trajectory Grading

Outcome grading asks whether the world ended up in the right state, such as the correct refund row, passing unit tests, or the right file diff. Trajectory grading asks whether the sequence of steps was acceptable, such as policy compliance, tool use, and efficiency. Use outcomes as the primary pass criterion and trajectory checks as targeted constraints, because over-specified paths punish valid solutions.

The two grading modes answer different questions, and conflating them is the most common design error in agent evals. Anthropic’s guidance states the principle bluntly: grade what the agent produced, not the path it took, so that valid creative approaches are not penalized. That is the right default. But it is not the whole story, and the exceptions are where trajectory evaluation earns its keep.

Why outcomes come first

Outcome graders are objective and cheap. A state diff on the database, a test suite exit code, or a schema validation either passes or fails, and it does so identically on every run. They are also robust to the thing agents do best: finding a route you did not anticipate. If your task is “cancel the order and refund the card”, then checking the orders table and the payments ledger is more reliable than checking that the agent called cancel_order before issue_refund, because a better agent may legitimately batch or reorder.

This is the design philosophy behind tau-bench. Its evaluation compares the final database state against an annotated goal state, so it does not need to know the dialogue that produced it. That choice is what lets the benchmark scale to many tasks without per-task human grading, and it is a pattern worth copying for your own domain whenever your environment has queryable state.

When trajectory checks are mandatory

An outcome can be right for the wrong reasons, and some wrong paths are unacceptable even when the end state is fine. Four categories justify explicit trajectory checks.

First, policy and safety constraints. An agent that refunds the correct order but first reads another customer’s record has reached the right outcome through a privacy violation. Rule-based checks over the trace, such as “no read of records outside the authenticated account”, catch this. For adversarial variants of the problem, see our write-up on agentic AI security and prompt injection, where trajectory assertions double as security tests.

Second, irreversible actions. Sending an email, wiring money, or deleting a resource leaves no state you can roll back in the real world. Assert that these calls occur at most once, only after required confirmations, and never in exploratory steps.

Third, efficiency. Two agents may both book the flight, but one used 6 tool calls and the other 41, with a different cost and latency profile. Track steps, tokens, and wall-clock time as secondary metrics with budgets, not as pass criteria.

Fourth, required evidence. In research or diagnosis tasks, the answer may be correct by luck. A check that the agent actually retrieved the relevant document before asserting a conclusion distinguishes grounded reasoning from a lucky guess. Our deeper treatment of these patterns lives in agent trajectory evaluation patterns.

Agent trajectory evaluation layers combining outcome graders, rule checks, and a calibrated LLM judge

Figure 2: Grader layers over a single trace and final state, feeding one task verdict and the pass^k computation.

Figure 2 shows the layering. The same trace and snapshot feed three kinds of grader. Outcome graders and rule-based trajectory checks are deterministic. The LLM judge handles what rules cannot express, and its output is trusted only to the degree it agrees with human labels. All three combine into a task verdict, which is then aggregated across trials.

A practical rule for choosing

Write the outcome grader first. Run the agent fifty times and read the failures. Each failure that the outcome grader scored as pass but that a human finds unacceptable becomes a candidate trajectory rule. Each failure scored as fail that a human finds acceptable means the outcome grader is too strict and must be loosened. This loop converges quickly, and it keeps your trajectory rules grounded in observed behavior rather than speculation about what an agent might do.

A subtle warning: exact-match trajectory comparison, where the agent’s tool-call sequence must equal a reference sequence, looks rigorous and is brittle. Anything that is order-insensitive in reality should be graded as a set or by state. Reserve strict ordering for genuine dependencies, such as authenticate before mutate. Partial credit is a better fit for multi-step tasks than a hard zero, because it shows progress during development; keep a strict binary pass for the CI gate.

Simulated Environments, Users, and LLM Judges

Repeatable agent trials need three simulated pieces: a tool sandbox whose state resets from a seed, a simulated user for conversational tasks, and an LLM judge for qualities that code cannot check. Each is a source of variance and bias you must control, because a noisy harness hides real regressions.

The tool sandbox

The sandbox is the agent’s world. Its job is to make tool calls deterministic enough that a difference in score reflects a difference in the agent. Practical techniques follow.

Seed the state from a file and reset it per trial. For a retail-style agent this is a small database with customers, orders, and inventory. For a coding agent it is a repository checkout at a known commit. For an operations agent it may be a fake cluster, which is the approach in our agentic SRE incident-response architecture, where replayed telemetry stands in for production.

Mock external services at the network boundary and record the responses. A tool that calls a live weather API or a live search engine returns different data on each run, and your trial-to-trial variance now includes the internet. Where you must call something live, pin the response with a recorded fixture and run a small separate “live” suite on a schedule.

Keep tool schemas identical to production. The most expensive eval bug is a sandbox whose tools differ from the real ones, so the agent learns, or is tuned against, a world that does not exist. Generate the sandbox tool definitions from the same source as the production ones.

The simulated user

Conversational agents cannot be tested with a single prompt, because the interesting behavior emerges over turns: the agent asks a clarifying question, the user answers incompletely, the agent must recover. tau-bench’s answer is to simulate the user with a language model given a persona and a hidden goal, and the later tau2 work extends this to a user who can act on the shared environment as well. A simulated user turns a conversation into a repeatable experiment.

Sequence of a simulated-user evaluation trial between harness, simulated user, agent, and tool sandbox

Figure 3: One trial of a conversational agent eval, from state reset to final snapshot and grading.

Figure 3 shows the control flow. The harness resets state, hands the simulated user its persona and goal, and the conversation proceeds between the user and the agent while the agent calls tools. After the final answer, the harness snapshots state and runs graders. Note that the harness, not the agent, owns the clock and the reset.

The simulated user is itself a model, so it introduces its own failure modes, and these deserve attention.

  • Goal leakage. If the user prompt makes the goal too easy to infer, the agent succeeds without needing to ask good questions. Give the user only the information a real person would volunteer.
  • Over-helpfulness. LLM users tend to cooperate more smoothly than real customers. Add personas that are terse, change their mind, or supply information out of order.
  • Drift from script. A user model may invent facts that contradict the scenario, creating failures that are not the agent’s fault. Check a sample of transcripts and add a guard that flags contradictions with the scenario record.
  • Coupled variance. Use a different, fixed model version for the user than the one under test, and pin its temperature. Otherwise upgrading the agent silently changes the user.

For high-value tasks, supplement simulated users with a smaller set of scripted turns, where the user’s messages are fixed. Scripted tasks have zero user variance and make excellent regression anchors, while simulated tasks provide breadth.

LLM judges for agents

Some qualities are real and unmeasurable by code: whether an explanation was accurate and appropriately hedged, whether a refusal was polite, whether a summary omitted a critical caveat. An LLM judge scores these against a rubric. The foundational evidence is encouraging: Zheng et al. found that a strong judge such as GPT-4 achieved over 80% agreement with human preferences, comparable to the agreement among humans themselves, while also documenting position, verbosity, and self-enhancement biases and limited reasoning ability (arXiv 2306.05685). Note that this study concerned chat answers, not long agent transcripts, so treat the 80% figure as a sanity anchor rather than a guarantee for your domain. Our pipeline-level guide to LLM-as-judge evaluation covers the mechanics in depth; here are the agent-specific rules.

Grade one dimension per call. A judge asked to rate “overall quality” blends criteria and becomes hard to calibrate. Ask separate questions: did the agent state a refund amount that matches the tool result; did it disclose the fee; did it stay in scope.

Judge the evidence, not the vibe. Give the judge the relevant tool outputs along with the claim. A judge that reads only the final message cannot detect that a stated order number was invented. Hallucinated tool results are best caught by comparing claims to the trace.

Randomize and control for bias. For pairwise comparisons, swap position and require agreement across both orders. Prefer a different model family for the judge than the agent to reduce self-enhancement. Cap length effects by instructing the rubric to ignore verbosity and by checking that scores do not correlate with token count.

Give the judge a way out. Include an “unable to determine” option so ambiguous cases route to human review instead of producing a coin flip dressed as a score.

Calibrate against humans, and keep calibrating. Label a few dozen transcripts by hand, compute agreement with the judge using Cohen’s kappa or simple percent agreement per rubric item, and track it over time. A judge model update is a change to your measurement instrument. Keep a frozen human-labeled set and rerun it whenever the judge model or prompt changes, which is why Figure 4 includes a judge-drift check.

The healthy division of labor is: code for facts, rules for policy, a judge for nuance, humans for calibration. If a quality can be checked with code, do not use a judge. It is slower, costlier, and noisier for no gain.

The Statistics of Small Suites: pass@k, pass^k, and Confidence

Agents are stochastic, so a single run per task proves little. Report both pass@k, the chance that at least one of k attempts succeeds, and pass^k, the chance that all k succeed. They answer opposite questions: pass@k measures capability headroom, while pass^k measures the reliability a customer actually experiences.

The difference is large. Anthropic’s guidance notes that as attempts increase, pass@k rises while pass^k falls. Consider an agent with a true per-trial success probability of 0.8 on a task, assuming independent trials. Then pass@1 is 0.80 and pass@4 is 1 minus 0.2 to the fourth power, which is about 0.998. But pass^4 is 0.8 to the fourth power, about 0.41, and pass^8 is about 0.17. Same agent, same task. If a user will hit that workflow repeatedly, the 0.17 describes their experience far better than the 0.80. This is exactly the pattern tau-bench’s authors highlighted when they reported pass^8 under 25% for an agent whose single-trial success was under 50% (arXiv 2406.12045). The 0.8 example above is illustrative arithmetic, not a measured result.

When you run n trials of a task and observe c successes, the standard unbiased estimator uses combinations, so you can estimate pass^k for any k up to n without rerunning:

from math import comb

def pass_hat_k(n: int, c: int, k: int) -> float:
    """Estimate P(all k sampled trials succeed) from c successes in n trials."""
    if k > n:
        raise ValueError("k cannot exceed n")
    return comb(c, k) / comb(n, k)

def pass_at_k(n: int, c: int, k: int) -> float:
    """Estimate P(at least one of k sampled trials succeeds)."""
    if n - c < k:
        return 1.0
    return 1.0 - comb(n - c, k) / comb(n, k)

def suite_score(results: dict[str, tuple[int, int]], k: int) -> float:
    """results maps task_id -> (n_trials, n_success). Mean pass^k across tasks."""
    scores = [pass_hat_k(n, c, k) for n, c in results.values()]
    return sum(scores) / len(scores)

# Example: 8 trials, 6 succeeded
print(pass_hat_k(8, 6, 1))  # 0.75
print(pass_hat_k(8, 6, 4))  # 15/70 = 0.214...
print(pass_at_k(8, 6, 4))   # 1.0

Small suites are noisy: do the interval math

A 50-task suite is a good start, and it is also statistically coarse. If the true pass rate is 0.8, the standard error of the observed rate on 50 independent tasks is the square root of 0.8 times 0.2 divided by 50, about 0.057. A 95% interval is therefore roughly plus or minus 11 percentage points. A change from 80% to 74% on that suite is inside the noise, and treating it as a regression, or as an improvement in the other direction, is a classic way to chase ghosts.

There are three honest responses. Run more trials per task to shrink within-task variance, which helps with stochastic agents but not with a small task set. Add tasks, which shrinks the dominant source of uncertainty. And compare paired results: because every commit runs the same tasks, use a paired test on per-task differences, or a bootstrap over tasks, which is much more sensitive than comparing two independent pass rates. Report the interval in the CI output so a reviewer sees “74% (68 to 80), baseline 80%” instead of a bare number.

Where variance comes from

Separate three sources. Task-level variance is that some tasks are simply harder, which is fixed across runs. Agent stochasticity is that sampling temperature and tool timing change the trajectory. Environment and judge variance is that flaky mocks and noisy judges add error. Decompose by running a few tasks many times: if the same task flips between pass and fail on identical inputs, you are looking at agent or environment noise. If it flips only across commits, you are looking at a real change. That experiment, cheap to run once, tells you how many trials your CI actually needs.

Wiring the Harness into CI and Production

A useful agent eval gate runs a fast smoke suite on every change, a full regression suite before merge, compares results to a stored baseline with a confidence interval, and feeds production failures back as new tasks. Speed, statistical honesty, and a closed feedback loop matter more than suite size.

Agent eval harness CI gate flow with smoke suite, regression suite, baseline comparison, judge drift check, and production feedback loop

Figure 4: CI gate and feedback loop for agent regression testing, from pull request to production traces and new tasks.

Tiered suites

Running everything on every commit is too slow and too expensive; running nothing until release is too late. Use tiers.

The smoke tier is 15 to 30 high-signal tasks with two or three trials each, finishing in minutes. Its job is to catch catastrophic breakage such as a malformed tool schema, a broken system prompt, or an authentication failure. Fail fast here.

The regression tier is the full suite with enough trials per task to estimate pass^k at the k your product needs. Run it on pull requests that touch prompts, tool definitions, model identifiers, or orchestration code, and nightly otherwise. Because inference cost scales with tasks times trials times steps, set a budget and let the harness report cost per run as a first-class metric.

The capability tier holds hard tasks with low pass rates. It does not gate merges. It tracks progress and helps you decide when a task has become reliable enough to graduate into the regression suite. Graduation is a useful ritual: a task moves when its pass^k stays above a threshold across several consecutive runs.

Gating rules that survive contact with noise

Gate on the paired difference against the baseline, not on an absolute number. A rule such as “block if suite pass^k drops by more than the margin with 95% confidence, or if any critical task regresses” balances sensitivity against false alarms. Critical tasks, those tied to safety, money, or legal exposure, should gate individually with no averaging, because a mean can hide a catastrophic task failure behind many easy wins.

Define a retry policy deliberately. Re-running a failed suite until it passes is a way to launder noise into false confidence. If you allow a rerun, require that the aggregate over all runs, not the best run, be reported. Also pin everything you can: model snapshot identifiers rather than floating aliases, temperature, random seeds for the simulator, and the judge version. When a vendor updates a model behind an alias, an unpinned harness will show a regression with no commit to blame.

Version the instrument

Store, with every result, the suite version and the grader version. When you fix a buggy grader, scores shift for reasons unrelated to the agent. A new grader version should trigger a re-baseline: rerun the previous commit with the new grader so the comparison stays like for like. The judge-drift check in Figure 4 is the same idea applied to the model-based graders, using the frozen human-labeled set.

Closing the loop with production

Offline suites are a model of reality, and reality changes. Sample production traces, run the same graders over them where outcomes are determinable, and route anomalies to human review. The reference point is the observability stack in our LLM observability architecture guide: because the evaluation and production traces share a schema, a graded production failure can be converted into a task record by capturing the starting state and the user goal. Each incident then becomes a permanent regression test. Over months this is how a suite acquires the long tail that no one could have designed up front.

For the production side, prefer cheap signals first. Tool-call error rates, retries, escalation to a human, and abandoned sessions are all available without a judge. Use judge-based sampling on a small fraction of traffic to detect quality drift, and keep the sampling rate low enough that the judge cost remains a rounding error on the agent’s own cost.

An example task record

A minimal task definition, shown as YAML for readability, makes the pieces concrete. The field names are illustrative, not a standard.

id: refund-partial-after-return-017
seed: seeds/retail_small.sqlite
user:
  persona: "terse, mildly impatient, gives order id only when asked"
  goal: "refund the damaged mug from order 4412 to the original card"
policy_refs: [refund_policy_v7]
graders:
  outcome:
    - sql: "SELECT status FROM refunds WHERE order_id=4412"
      expect: "issued"
    - sql: "SELECT COUNT(*) FROM refunds WHERE order_id=4412"
      expect: 1
  trajectory:
    - never_call: [read_customer_record_other_account]
    - must_precede: {first: verify_identity, then: issue_refund}
    - max_calls: {issue_refund: 1}
  judge:
    - rubric: "States the refund amount exactly as returned by the tool"
      evidence: [tool_result.issue_refund]
budgets: {max_steps: 25, max_tokens: 60000}
trials: 8

Notice how the pieces divide the labor. The SQL checks verify the world. The trajectory rules encode the three constraints that outcomes cannot see: identity checked first, no peeking at other accounts, exactly one refund. The judge item handles a claim about language, anchored to a specific tool result. And the budget turns efficiency into a measurable boundary.

Trade-offs, Gotchas, and What Goes Wrong

Harnesses fail in predictable ways, and most of them produce false confidence instead of visible errors. The worst outcome of an eval program is a green dashboard that does not correlate with production quality.

Overfitting to the suite. If engineers tune prompts until the 50 tasks pass, you have trained on the test set. Hold out a private subset that is not used during development, and refresh it from new production failures. Watch for a growing gap between suite scores and online metrics; that gap is the overfitting signal.

Benchmark gaming and contamination. Public benchmarks leak into training data, and some task sets turn out to contain flawed items. For example, community efforts have published corrected versions of tau2-bench tasks where definitions and expected actions did not align with the stated policies, a reminder that benchmark ground truth deserves audit, not worship. Treat your own tasks the same way: periodically review the failing tasks, because a task that no agent can pass is often a broken task.

Brittle graders. Code graders that match exact strings fail valid variation, and judges drift. Both produce regressions that are not regressions. The remedy is to read failure transcripts regularly. There is no substitute for a human reading twenty failed traces; automated scores tell you where to look, not what is true.

Reward hacking inside the sandbox. Agents optimize what is measured. A coding agent may edit the tests, a conversational agent may learn that a polite closing phrase satisfies the judge, an agent may exploit a sandbox quirk unavailable in production. Make graders read-only to the agent, grade outcomes from a separate trusted process, and keep a few canary tasks designed to detect shortcuts.

Cost and latency. A suite of 100 tasks times 8 trials times 30 steps is 24,000 model calls before you count judge calls. At any realistic price this adds up, and slow suites get skipped. Budget by tier, cache deterministic tool results where it does not defeat the test, and consider a cheaper model as simulated user and judge only after verifying agreement with the stronger one.

Non-determinism you cannot remove. Even at temperature zero, hosted models are not guaranteed bit-identical across runs. Plan for variance rather than assuming it away; that is the entire purpose of reporting intervals and pass^k.

Measuring the wrong thing. A high task-success score on tasks that do not reflect real user intents is worthless. Weight tasks by production frequency and by business impact, and revisit the weighting quarterly. Averaging across a suite built from convenient rather than representative tasks produces a number that moves but means little.

Simulator realism gap. A simulated user is more patient and articulate than most humans. Agents that ace the simulation can still struggle with real conversations, which is why the production feedback loop is not optional.

Practical Recommendations

Start small and make the loop real before making the suite large. A 30-task suite with isolation, trace capture, a pass^k report, and a CI gate is more valuable than a 500-task suite that runs quarterly. Build the pipeline end to end first, then widen the coverage.

Choose graders by the cheapest reliable method: state checks for facts, rules for policy and irreversible actions, a calibrated judge for language quality, humans for calibration and for reading failures. Make outcomes the pass criterion and trajectory rules the constraints, and keep the number of trajectory rules small and justified by observed failures.

Treat the harness as production software. Version tasks, graders, and judge prompts. Pin model snapshots. Record the cost of each run. Review it when it surprises you, because a harness bug that hides a regression is worse than no harness.

Finally, decide up front what reliability your product needs. A back-office assistant that a human double-checks may tolerate a pass^1 target, while an autonomous workflow that runs a hundred times a day needs a high pass^k at a k near its daily repetition. The target decides how many trials you run and where the gate sits.

Rollout checklist

  • Write 20 to 50 tasks from real or realistic failures; store them in git with seeds.
  • Reset the sandbox for every trial; verify two trials cannot share state.
  • Capture full traces in a schema you also use in production.
  • Write outcome graders first; add trajectory rules only for observed failure classes.
  • Add judge graders for one dimension each, with an “unable to determine” option.
  • Label 30 to 50 transcripts by hand; measure judge agreement; freeze that set.
  • Report pass@k and pass^k with intervals; run enough trials to estimate your k.
  • Gate on paired differences against a baseline; gate critical tasks individually.
  • Pin model snapshots, temperature, simulator seeds, and judge version.
  • Convert every production incident into a new task within the week.
  • Hold out a private subset; compare suite scores to online metrics monthly.

Frequently Asked Questions

What is the difference between trajectory evals and outcome evals for AI agents?

Outcome evals check the final state of the environment, such as a database row, a passing test suite, or a correct file diff. Trajectory evals check the path: which tools were called, in what order, with what arguments, and whether policies were respected. Use outcomes as the primary pass criterion because they allow valid alternative approaches, and add trajectory checks for irreversible actions, safety constraints, required evidence, and efficiency budgets.

What does pass^k mean and why does it matter for agents?

pass^k is the probability that an agent succeeds on all k independent trials of the same task, introduced for conversational agents by the tau-bench authors to measure reliability. It matters because a high single-trial success rate can hide poor consistency: with a per-trial success of 0.8, pass^8 is only about 0.17 under independence. Customers who repeat a workflow experience pass^k, not pass^1, so report both with confidence intervals.

How many tasks do I need for a useful agent eval suite?

Anthropic’s engineering guidance recommends starting with 20 to 50 realistic tasks drawn from actual failures instead of waiting for hundreds. Early changes tend to have large effects, so a small suite already discriminates. Be aware of the noise: on 50 tasks, an 80% pass rate has roughly a plus or minus 11 point 95% interval, so use paired comparisons against a baseline and grow the suite as incidents occur.

Can I trust an LLM as a judge for agent evaluation?

Partly. Research on chat evaluation found strong judges reaching over 80% agreement with human preferences, similar to human-to-human agreement, but with documented position, verbosity, and self-enhancement biases. For agents, grade one dimension per call, give the judge the tool evidence, randomize ordering, add an unable-to-determine option, and calibrate against a frozen human-labeled set. Use code or rules whenever a property can be checked deterministically.

What is a simulated user in agent evaluation?

A simulated user is a language model given a persona and a hidden goal that converses with the agent under test, making multi-turn tasks repeatable. tau-bench popularized the approach for tool-using customer-service agents. Its weaknesses are goal leakage, excessive cooperation, and invented facts, so pin the simulator model and temperature, vary personas, sample transcripts for contradictions, and complement it with a few fully scripted tasks.

How do I stop agent regressions from reaching production?

Run a fast smoke suite on every change and a fuller regression suite on changes to prompts, tools, or model versions. Gate on the paired difference from a stored baseline with a confidence interval, and gate critical tasks individually. Pin model snapshots and judge versions, then feed production failures back as new tasks. Evals reduce regressions rather than eliminate them, so keep production monitoring in place.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *