Key Takeaways
- Agent behavior changes from run to run, so a handful of successful demos cannot show that a prompt, model, or tool change actually made things better.
- The only way to know is to run a baseline and a candidate on the same golden dataset and compare them against release criteria set before the test starts.
- A continuous evaluation loop connects two stages that share the same rubric and evaluators: pre-production experiments that compare configurations before release, and production monitoring that scores live traces after release. Human judgment sits underneath both stages.
- People write the rubric and annotate real cases, which calibrates the deterministic checks and LLM-as-a-judge evaluators doing the scoring, and people review production failures so the ones that reproduce become regression cases the next candidate has to pass.
AI agents can produce convincing answers even when their claims are inaccurate or their actions lack required approval. For example, a checkout agent might claim an order was placed even though the checkout system rejected the request, or the agent may submit the order before the customer confirms the total. The agent's response to the user may not reflect the actual outcome or whether it followed the required approval process.
As agents handle longer conversations and more tools, these types of failures become harder to spot. To determine whether an agent succeeded, teams need evaluation criteria tailored to the task and trace-records from the interaction. For a checkout agent, the trace-records should show what the customer approved, what the agent attempted, what changed in the underlying system, and what the agent finally conveyed. Continually evaluating this information helps teams find failures, test fixes, and improve the agent's behavior over time.
A continuous evaluation loop connects pre-production experiments with trace-records from live interactions to guide and verify iterative improvements to an agent. This blog explains how to define rubrics and evaluators, compare baseline and candidate configurations on a golden dataset, monitor production behavior, and turn production failures into regression cases for evaluating future changes.
Throughout this blog, we use a retail checkout agent as an example. The agent helps a customer find an item, checks inventory, adds the item to a cart, presents the final order and total, waits for the customer to approve the purchase, submits the order, and reports the result.
.jpg)
1. What Are Evals, and Why Do AI Agents Need Them?
An evaluation, or eval, is a structured assessment of whether an agent handled a task as intended. It defines the task and its context, states what counts as success or failure, and checks trace-records of the agent's actions and outcome against those criteria. Teams can run the task in a defined environment before release or apply the criteria to a recorded production interaction. For a retail checkout agent, a pre-production case might ask it to find the requested item, confirm it is in stock, add it to a cart, show the total, wait for confirmation, and then submit the order. The expected outcome includes both the completed order and the conditions or sequence in which it was placed.
A conventional API follows programmed logic, so the same request and system state generally produce the same result. A retail checkout agent, for example, interprets the customer's wording, selects tools and arguments, decides whether to submit the order after seeing tool results, and composes a response. The LLM may select different tools, arguments, actions, or responses for the same request, while changes in inventory, price, or checkout status can alter the information the agent receives along the way. A small change early in the interaction can therefore lead to a different cart action, order outcome, or final answer.
Because agent behavior varies, a few successful interactions cannot establish whether an update to the agent improved its behavior. Updating the system prompt, model, tool definitions, supplied context, or rules for taking action may improve one task while degrading another. Teams can run the same tasks repeatedly with the current agent and with the updated candidate agent to compare their behavior and outcomes. Running these comparisons as experiments shows whether the candidate meets the defined success criteria more often across cases and repeated trials, and whether the update introduces new failures.
2. Components of a Continuous Evaluation Loop
A continuous evaluation loop helps answer questions such as “Did the agent handle this task correctly?” into repeatable processes before and after release. Both use the same evaluation objectives, rubrics, evaluators, and trace records. In preproduction, a harness runs defined cases under controlled conditions so teams can compare configurations and make a release decision. In production, the trace record comes from live interactions, and an evaluation pipeline applies evaluators to eligible traces. Evaluation can run continuously, on a representative or risk based sample, or when another signal triggers it. These production results can reveal failures and unfamiliar behavior that the controlled dataset did not anticipate.
The table below defines the components that support such an evaluation loop.

3. Why Continuous Evaluation Requires Human Judgment
Automated evaluators can apply the same criteria repeatedly, but people must first decide what acceptable behavior means and which trace records support that judgment. They express those decisions in a rubric, which lists the criteria and judgments that can apply across cases.
Consider a reviewer examining a retail checkout trace from an early development run. The trace includes the customer's messages, the displayed proposal, and the order-submission event. The table below shows how a checkout rubric turns the desired behavior into judgements.
Case rule: The agent must obtain the customer's approval before submitting the order. If it does not, the case fails. If the agent obtained approval, the case passes only when every other criterion that applies to the case also passes. A reviewer marks a criterion not applicable only when the rubric says it does not apply. If no criterion fails but the trace record is insufficient for one or more applicable criteria, the overall result is inconclusive and requires human review.
The reviewer applies the rubric and records a human annotation with the judgment, supporting trace records, and rationale. Once reviewed, the case and its annotation can be added to the golden dataset. The annotation provides the reference judgment used to test and calibrate software-based evaluators that apply the rubric without requiring a person to review every trace. These evaluators may use deterministic rules or an LLM-as-a-judge evaluator.
Example Human Annotation
Overall judgment: Fail because approval compliance is a critical criterion.
The golden dataset provides reviewed human annotations for testing and calibrating evaluators. Each rubric criterion should use an evaluator suited to its outcome. A deterministic check can verify conditions, such as whether an order record exists or whether approval occurred before submission. An LLM-as-a-judge evaluator can assess semantic criteria, such as whether the agent's final response accurately explains the events in the trace.
Teams calibrate an evaluator by comparing its judgments with human annotations and checking disagreements against the supporting trace evidence. This review may reveal an unclear rubric criterion, an evaluator error, or an annotation that needs correction. LLM-as-a-judge evaluators should be tested against high-quality examples with reviewed human judgments. Adding reviewed edge cases to the golden dataset helps test how the evaluator handles similar cases in future runs. Versioning the rubric, evaluators, annotations, and dataset helps teams interpret results and understand what changed between evaluations [1], [2].
4. Running Controlled Pre-production Experiments
Pre-production experiments provide a controlled way to test an agent update before it reaches users. The golden dataset, rubric, and evaluators support a comparison between the current and proposed configurations, showing whether the candidate meets the predefined release criteria without introducing unacceptable failures.
A controlled preproduction experiment begins by splitting selected cases from the golden dataset into development and held-out test sets. The development cases guide refinements to the prompt, model, tools, or agent logic. Keeping the held-out test cases out of the tuning process reduces the risk of optimizing the agent for familiar examples and tests whether the improvement extends to cases that were not used during development. Release criteria are set in advance to avoid adapting it to the result of the experiment. For example, the checkout-agent candidate might be released only if it never submits an order without approval and is at least as reliable as the baseline at submitting orders with the correct items, quantities, and totals.
The baseline, representing the current configuration, and the candidate, representing the proposed configuration, run on the same cases. Each configuration starts from the same initial state, receives equivalent inputs, and runs under the same operating conditions. Cases can be repeated to measure variation. To keep the experiment reproducible, its record includes the versions of the agent, dataset, rubric, and evaluators.
Each trial produces a trace that deterministic checks and LLM-as-a-judge evaluators score against the rubric before the pipeline aggregates the results and retains links to the supporting trace-records. Human reviewers examine critical failures, evaluator disagreements, inconclusive results, and a sample of passes. They confirm or correct the judgments and identify whether the problem lies with the agent, a dependency, the case, the rubric, or the evaluator.
The release decision must consider aggregate performance alongside individual failures. Even when a candidate improves overall, it should not proceed to release if it introduces a failure in a critical workflow. Changes to the prompt, model, tools, or agent logic create a new candidate for testing, so fresh trials are required to observe how the revised configuration behaves.
Pre-production evaluation may also reveal that the rubric or evaluator needs to be updated:
- Human review may expose unclear or missing rubric criteria, while disagreements with human annotations may show that an evaluator needs recalibration.
- A revised rubric should be reviewed against annotated cases to confirm that its criteria are clear and can be applied consistently.
- A revised evaluator should be recalibrated against reviewed annotations to confirm that it applies the rubric reliably. If a rubric change affects an automated evaluator, that evaluator must also be updated and recalibrated.
A successful pre-production experiment provides a degree of confidence that a candidate agent is ready for release under known, controlled conditions, but it cannot guarantee how the agent will behave in production. We will explore production monitoring next.
5. Monitoring AI Agents in Production
Production monitoring shows how the agent behaves under live conditions and whether the results observed in preproduction carry over when it encounters customer requests not covered by the test cases. It also reveals how the agent responds to changes in its operating environment, such as inventory or price updates and failures in services such as the catalog or checkout API.
Production evaluation depends on a complete trace of each interaction. A correlated trace uses shared identifiers to link events from the same customer request, such as the customer's message, the model response, an inventory lookup, the order-submission call, the API result, and the agent's final reply. Reviewers can then determine whether a problem came from the agent, stale data, or a failed service. Traces should preserve the context needed for review while redacting sensitive information and limiting access and retention.
Operational monitoring, behavioral evaluation, runtime controls, and human review each examine a different aspect of production behavior and support a different type of objective.
Operational monitoring shows whether the deployed system remains responsive and reliable. In addition to Google SRE's four golden signals [3] for performance monitoring (latency, traffic, errors, and saturation), agentic AI systems should track token use and cost per interaction. In an agentic AI system, traffic can also be measured by the rate of customer sessions and tool calls, and signs of saturation can include growing request backlogs or frequent failures caused by model or API rate limits. These signals can reveal problems that behavioral evaluation may miss because an agent may eventually produce the correct answer after repeated tool retries even though the added latency and cost degrade the customer experience.
Behavioral evaluation assesses whether the agent completed the task correctly and followed the required process. In the checkout example, it examines whether the agent submitted the correct order, obtained the customer's approval before submission, and accurately reported the verified outcome in its final response. When a production interaction involves a task covered by an existing rubric, deterministic checks and LLM-as-a-judge evaluators can score its trace. If the interaction includes behavior that the rubric does not cover, a reviewer must examine the trace and decide how it should be judged. For example, the rubric may not specify whether the agent must request approval again when the order total changes after the customer has approved it. Reviewing such cases helps determine the appropriate judgment and whether the rubric should be updated to cover the behavior. Operational metrics and behavioral results should link to the same trace so reviewers can inspect the trace-records behind a signal or score.
Runtime controls and human review complement monitoring and behavioral evaluation at different stages. Runtime controls enforce requirements as the agent acts, such as blocking submit_order until the customer approves the order and total. Human reviewers examine completed traces and evaluator results to confirm judgments, diagnose failures, and identify gaps in the rubric or evaluator.
6. Turning Production Evaluation into Continuous Improvement
It is not practical to review every production interaction manually. Automated evaluations can assess traces at scale, and their results, together with other monitoring signals, help select cases for human review. These signals can include:
- failed or inconclusive evaluation results
- inaccurate final responses
- customer corrections
- negative feedback
- unusual numbers of tool calls
- attempts to take actions that are not permitted
Targeted selection can distort the general picture unless it is balanced with a representative sample of ordinary interactions, including successful refusals, clarifying questions, and recoveries after tool errors. For example, repeated corrections about prices or stock status can form a review cohort that is compared with ordinary checkout conversations to show how often the issue occurs and what successful behavior looks like.
Each selected interaction should retain its trace reference, agent configuration, relevant inputs and system state, and reason for selection. Near-identical cases should be deduplicated so that a single incident is not overrepresented in the evaluation dataset. Before production trace-records are added to a dataset, customer identifiers should be removed or replaced, and the resulting dataset should use the same access controls and retention periods as the source data.
Selecting a production interaction sends it to human review. Reviewers verify its trace and record a human annotation, which can be used to validate LLM-as-a-judge evaluators. If the failure can be reproduced, reviewers define its inputs, initial state, and expected behavior, then add the resulting regression case to the golden dataset. For example, if the agent recommends an item listed in the catalog even though inventory reports it as out of stock, the regression case preserves the customer's request and recreates the catalog response that listed the item and the inventory response that reported it as out of stock. The expected behavior requires the agent to check inventory before listing the item as available. The saved trace provides provenance of the past behavior and the regression case tests future agent versions.
Reviewed findings can be prioritized by impact, frequency, and confidence in the diagnosis. A rare order placed without customer approval may deserve attention more than a common wording problem because its consequences are more serious. Reviewers use the trace to determine why the failure occurred and whether the fix belongs in the agent, a data source, or another dependency. In the out-of-stock example, the appropriate fix depends on the cause. If the agent failed to check inventory, its logic should be updated. If stale data or an unclear tool response caused the recommendation, the corresponding data source or tool should be fixed.
The new regression case is included in the next controlled comparison between the baseline and candidate agent configuration. The candidate should pass the new case, preserve correct behavior on ordinary checkout cases, and meet the other release criteria. After release, production monitoring shows whether the improvement continues to work under live conditions. New failures are selected, reviewed, and annotated; when they can be reproduced, they become regression cases or lead to updates to the rubric or evaluators. This cycle produces a loop where production evaluation results are sent to the next round of engineering changes, testing, and release decisions.
7. How Fiddler Supports the Continuous Evaluation Loop
The continuous evaluation loop is tool agnostic, but implementing it requires capabilities that span pre-production experiments, production monitoring, human review, and runtime controls. Fiddler supports these stages through its evaluation, experimentation, observability, and policy enforcement capabilities.
Fiddler connects production signals to the traces, annotations, evaluations, and experiments used in the next improvement cycle.
8. Where to Start
Teams do not need a scaled evaluation program to begin. Start with one high-impact task and a critical failure. For the checkout agent, that might be submitting an order without approval or claiming that a rejected order succeeded. Begin by defining the task's success conditions, writing a focused rubric, and curating a small golden dataset from reviewed cases. Next, establish a process for selecting production trace-records, sending it to human review, and incorporating useful cases into future experiments.
Expand the loop as the agent takes on more responsibilities. Over time, it creates a repeatable process in which proposed changes are tested before release and production failures improve the tests used for the next version.
Ready to put a continuous evaluation loop around your own agents? Request a demo to see how Fiddler connects pre-production experiments, production monitoring, human review, and runtime controls in one platform.
References
[1] M. Grace, J. Hadfield, R. Olivares, and J. De Jonghe, “Demystifying evals for AI agents,” Anthropic, Jan. 9, 2026. [Online]. Available: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
[2] OpenAI, “Graders,” OpenAI API Documentation. [Online]. Available: https://developers.openai.com/api/docs/guides/graders (accessed Oct. 6, 2026).
[3] R. Ewaschuk, “Monitoring distributed systems,” in Site Reliability Engineering: How Google Runs Production Systems, B. Beyer, Ed. Sebastopol, CA, USA: O’Reilly Media, 2017, ch. 6. [Online]. Available: https://sre.google/sre-book/monitoring-distributed-systems/

