Key Takeaways
- An eval pairs each test case with a grader that returns a score, so you can compare a prompt, model, or configuration change against a baseline before and after release.
- Choose scorers in order of cost: start with deterministic code-based graders, add LLM-as-a-Judge only when the failure is semantic, and reserve human review for calibration labels and disputed cases.
- Validate every judge against human labels with Cohen's kappa, and treat 0.80 or higher on 100 or more labeled cases as ship-ready, since raw agreement can hide weak alignment.
- Agent evals score the full trajectory, including tool choice, arguments, repeats, memory, and end state, which requires span-level telemetry.
- Tie each metric to one threshold that gates releases in CI and fires the matching production alert, and use public benchmarks only to break ties between candidate models.
AI evaluation is the systematic, repeatable testing of a model, LLM application, or agent against a curated dataset with defined scoring criteria. Each test case pairs an input with a grader that returns a score, so you can compare a prompt, model, or configuration change against a baseline before and after release.
Production AI fails in ways deterministic tests cannot catch. A prompt edit that trims tokens can pass every schema check yet cause a retrieval agent to answer from stale context or fabricate a tool argument. It may also loop until the token budget is gone. AI evaluation gives you a scored record of behavior that you can gate a release on and use for production alerts. You can also retain that record as audit evidence for each release decision.
What Is AI Evaluation
An eval measures six properties of an AI system: output quality, factual accuracy, safety, cost, latency, and, for agents, the trajectory of intermediate steps. Three kinds of scorer produce those measurements:
A unit test asserts that a function returns one value and fails otherwise. An LLM returns a distribution of plausible outputs, so a faithful answer on one run and a fabricated citation on the next are both normal behavior for the same prompt. Evals therefore report graded results and compare them against a threshold. Anthropic recommends that capability evals start with a low pass rate to leave room for improvement, while regression evals should hold near 100 percent [1].
Where AI Evaluation Fits: Classical ML, Generative, and Agentic
Classical ML evaluation scores predictions against labels with precision, recall, F1, and AUC. Generative evaluation has no single correct output, so it scores faithfulness, groundedness, relevance, toxicity, and rubric-based quality. Agentic evaluation adds the path: tool choice, order, arguments, and whether the final state matches the goal.
Classical binary metrics still matter in one place. True and false positive rates measure a binary judge or guardrail against human labels more precisely than raw agreement.
How AI Outputs Get Scored
Most production systems use all three scoring methods at different points in the pipeline.
Code-Based Graders
Exact match, regex, JSON schema validation, and numeric bounds suffice when output has a structural contract. A tool call must parse as JSON with a known argument set, and a refund amount must not exceed the order total. Write these checks first; add a model-based judge only when the failure is semantic.
LLM-as-a-Judge
A judge reads the input, candidate output, any retrieved context or reference answer, and a rubric, then returns a label or score. Anthropic recommends one judge per dimension with a structured rubric, calibrated against human graders [1].
Judges carry documented position bias. Under a default prompt, GPT-4 returned the same pairwise verdict regardless of answer order 65 percent of the time and Claude-v1 23.8 percent [2]. Run the judge twice with swapped order and accept a verdict only when both runs agree [2]. A provider update to the model behind the API can change judge behavior, so pin the judge version and re-validate on every change.
Human Review
Human review covers nuance, expert domains such as clinical or legal content, and the labeled set you calibrate judges against. Human throughput is orders of magnitude below automated scoring, which rules out reviewing production traffic. Reserve humans for a 100–200 case calibration set and borderline cases the judge flags. Use periodic audits to check judge output.
What to Measure: Quality, Safety, and Efficiency Metrics
Metrics fall into three buckets:
Attach each score to a threshold that acts as a release gate before deployment and an alert or redaction rule in production.
Offline vs Online Evaluation
Offline runs a curated golden set with reference outputs on every commit and nightly to catch regressions before release. Online scores live traffic, which usually lacks reference outputs, with reference-free scorers such as faithfulness to retrieved context, toxicity, and PII/PHI detection. The evaluation team adds each failure found online as a new offline case. Use the same threshold per metric in both loops.
Evaluating AI Agents
Where a single-turn eval scores one response, an agent eval scores a trajectory: the plan, each tool call and its arguments, memory reads, sub-agent handoffs, and the final environment state. In the MAST study of 1,642 annotated traces across seven agent frameworks, step repetition accounted for 15.7 percent of failures and failure to recognize termination conditions 12.4 percent [3].
Watch for four failure modes:
- Hallucinated tool calls: The agent invokes a tool that is not in its registry, or issues a redundant call that changes state twice.
- Wrong tool arguments: The tool exists but receives fabricated parameter values or malformed JSON.
- Infinite loops: The same action returns the same observation until the token budget or a step limit ends the run. Count repeats per trace and alert on them.
- Stale memory: The agent reasons from context that has since changed.
Score trajectories two ways. Compare the tool-call sequence against a reference in exact, in-order, or any-order mode, and grade the outcome. Anthropic warns against rigidly checking exact tool sequences because it penalizes valid alternative approaches [1]. Both require span-level telemetry: tool name, call ID, arguments, result, and parent span.
Building Your First Eval Set
Anthropic advises that “20–50 simple tasks drawn from real failures is a great start” [1].
- Pull 20–50 cases from production traces.
- Stratify across common intents and known edge cases. Include adversarial inputs.
- Add a case for every production incident and new intent.
- Keep the regression subset near 100 percent pass and expect the capability subset to fail often.
A worked example: a refund agent that reads an order and a policy excerpt alongside a customer message, then returns a decision. Twenty cases from support logs include 12 standard refunds and 5 edge cases (partial refunds, expired windows). The remaining 3 contain injected instructions in the order notes.
The code-based check validates the structural contract:
def check_refund(out, order):
return (
{"decision", "amount", "policy_clause"} <= out.keys()
and out["decision"] in {"approve", "deny", "escalate"}
and 0 <= out["amount"] <= order.total
)The judge scores faithfulness to policy:
Grade the refund agent's response for faithfulness to the policy excerpt.
Policy: {policy_excerpt}
Response: {response}
Every claim about eligibility or amount must be supported by the policy text.
Label High (all claims supported), Medium (one unsupported claim),
or Low (contradicts the policy or cites a clause that does not exist).
Return the label and the unsupported claim, if any.The following hypothetical scorecard for the 20 cases shows how thresholds become decisions:
Can You Trust the Judge
Validate every judge against human labels with a chance-corrected statistic, because raw agreement inflates on imbalanced data. Zheng et al. reported GPT-4 agreeing with human experts 85 percent of the time on MT-Bench, against 81 percent human–human agreement [2]. In Eugene Yan's comparison, Llama-3-8B reached 80 percent agreement with humans but a Cohen's κ of only 0.62, while GPT-4 reached κ 0.84 and the human–human κ was 0.97 [4].
Agreement varies by task. Bavaresco et al. measured GPT-4o's average κ at 0.28 ±0.32, from 0.84 on one dataset to 0.01 on medical safety [5]. Under the Landis and Koch scale, 0.61–0.80 is substantial agreement and 0.81–1.00 almost perfect, though the authors call the divisions arbitrary [6].
Recommended use by agreement level:
Re-run validation whenever the judge model version changes.
The Cost and Latency of Evals at Scale
The AlpacaEval README documents judge cost at $13.60 per 1,000 examples with GPT-4 and $5.50 per 1,000 with GPT-4 Turbo returning a single token [7]. Daily judge cost is responses per day × scored dimensions × per-call price, and current prices differ from those historical figures. That arithmetic drives the standard cadence. Run deterministic checks on every commit and judge suites nightly or pre-release.
Sampling production traffic at 5 or 10 percent cuts the bill, but rare events such as a PII/PHI leak may be missed. Fiddler frames evaluation total cost of ownership as three components:
Estimate all three for your traffic with the Evaluation TCO Calculator at fiddler.ai/evals-tco-calculator.
Public Benchmarks vs Your Own Evals
Public leaderboards such as MMLU and SWE-bench answer one question: how a base model compares with other base models on a fixed task.
Label errors weaken the signal, with Gema et al. estimating that 6.49 percent of MMLU questions contain errors [8]. Contamination compounds the problem. With answer options masked, ChatGPT reproduced the original MMLU option text 52 percent of the time and GPT-4 57 percent [9].
On February 23, 2026, OpenAI audited 138 hard SWE-bench Verified problems and found material issues in 59.4 percent of them [10]. The audit also showed that frontier models could reproduce the gold patch or verbatim problem-statement details on certain tasks, so OpenAI stopped reporting the score [10].
A leaderboard number helps when you are choosing among base models for a task the benchmark resembles and the gap between candidates exceeds the benchmark's known error rate. As release-review evidence that your application works, it is theater. Benchmark tasks rarely match your application's inputs. Build the custom eval set described above, run it against each candidate model, and let the benchmark break ties.
AI Evaluation Tools and Platforms
Tooling falls into three categories; judge each by whether a scored threshold can fail a CI build and fire the same production alert:
A common ownership pattern splits the work three ways: the application developer owns the case set and code-based graders, the domain expert owns the rubric and human labels, and the platform team owns scoring infrastructure and the thresholds linking CI to production alerts.
Fiddler AI Observability and Security Platform covers both the offline loop (pre-deployment evaluation on a curated golden set) and the online loop (production evaluation of live traffic) from one scoring layer. Fiddler evaluates 100+ metrics, and for evals it supports custom evaluators and comparisons across prompts, models, and configurations. Evaluation thresholds can connect pre-deployment tests to production alerts and enforcement rules, so the same score that blocks a release also triggers the production response.
Batteries-included Fiddler Centor Models score prompts and responses in-environment. Out-of-the-box and customizable models keep data in your environment and evaluate 100 percent of prompts and responses by default, at no per-evaluation cost, while processing millions of requests per day. Teams that want an external judge can route scoring to OpenAI, Anthropic, or Google through Bring Your Own Judge.
Request a demo to review how the required controls would fit your agent architecture and deployment environment.
References
[1] Anthropic, Demystifying Evals for AI Agents, https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
[2] Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023), https://arxiv.org/html/2306.05685v4
[3] Cemri et al., Why Do Multi-Agent LLM Systems Fail? (MAST, NeurIPS 2025), https://arxiv.org/html/2503.13657v3
[4] Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators, https://eugeneyan.com/writing/llm-evaluators/
[5] Bavaresco et al., LLMs Instead of Human Judges? A Large Scale Empirical Study Across 20 NLP Evaluation Tasks (ACL 2025), https://arxiv.org/html/2406.18403v2
[6] Landis and Koch, The Measurement of Observer Agreement for Categorical Data (1977), https://jjcurtin.github.io/book_iaml/pdfs/landis_1977_kappa.pdf
[7] AlpacaEval README, https://github.com/tatsu-lab/alpaca_eval/blob/main/README.md
[8] Gema et al., Are We Done with MMLU? (MMLU-Redux), https://arxiv.org/html/2406.04127v3
[9] Deng et al., Investigating Data Contamination in Modern Benchmarks for Large Language Models (NAACL 2024), https://aclanthology.org/2024.naacl-long.482/
[10] OpenAI, Why We No Longer Evaluate SWE-bench Verified, https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
Frequently Asked Questions
What Does an AI Evaluator Do?
It scores an output against one criterion using a deterministic function, an LLM judge with a rubric, or a purpose-built scoring model.
What Is the Difference Between Offline and Online Evaluation?
Offline runs a curated set with references before release; online scores live traffic without references and feeds failures back as new cases.
How Many Test Cases Should I Start With?
Start with 20–50 tasks from real failures [1] and hold 100 or more human-labeled cases for judge validation.
How Does Agent Evaluation Differ from Single-Turn LLM Evaluation?
Agent evaluation scores the full trajectory of tool choices, arguments, repeats, memory, and end state, not one response.
Which Metrics Should I Track?
Track faithfulness and task completion for quality. For safety, measure toxicity, PII/PHI, and injection. Track latency and cost per session for efficiency, with each metric tied to a threshold.
