Benchmark
A benchmark is a defined set of tasks and scoring conventions used to compare performance. An AI benchmark measures selected capabilities under its particular test conditions.
[Hendrycks et al.]In practice · hypothetical example
A team compares two models using the same subject questions and answer-scoring rules.
[Hendrycks et al.]A little deeper
MMLU, for example, tests knowledge across many subjects. Performance on a benchmark should be interpreted in terms of the tested tasks, not as a complete description of a system. [Hendrycks et al.]
A common mix-up
One benchmark score describes every real-world capability.
A benchmark covers its selected tasks and conditions. [Hendrycks et al.]
What makes two benchmark scores meaningfully comparable?
Sources & editorial notes
Evidence: supported. Primary-source support for this scoped entry; publication approved by the project owner.
- Measuring Massive Multitask Language Understanding ↗ (opens in new tab)Hendrycks et al. · Publication date unknown
Relevant section: Abstract
Last editorial review: 2026-09-13 by project-owner.
First observed in this corpus: Unknown.
Revision history
Revision 2 · Created 2026-09-13 · Updated 2026-09-13
Project owner approved the current content for publication. Existing evidence scope and limitations remain applicable.