Evaluation
Also called Eval, Evals
Evaluation is checking an AI system against defined tasks and success criteria. An evaluation may score an answer, inspect actions, or test whether the resulting environment matches the intended outcome.
[Anthropic]In practice · hypothetical example
A team gives a booking assistant ten test requests and checks whether each resulting reservation follows the requested dates and constraints.
[Anthropic]A little deeper
A task specifies the goal; a trial is one attempt; a grader scores a result. Repeated trials help reveal variability that a single successful demonstration can hide. [Anthropic]
A common mix-up
One successful demo proves the agent is dependable.
Success on one attempt does not establish performance across tasks or repeated trials. [Anthropic]
Why repeat an evaluation task?
Sources & editorial notes
Evidence: supported. Primary-source support for this scoped entry; publication approved by the project owner.
- Demystifying evals for AI agents ↗ (opens in new tab)Anthropic · Publication date unknown
Relevant section: The structure of an evaluation
Last editorial review: 2026-09-13 by project-owner.
First observed in this corpus: Unknown.
Revision history
Revision 2 · Created 2026-09-13 · Updated 2026-09-13
Project owner approved the current content for publication. Existing evidence scope and limitations remain applicable.