Benchmark (AI evaluation)

From ALT-TEXT
Revision as of 10:36, 7 September 2026 by imported>ALT-TEXT (Import: AI terminology and people glossary)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Benchmark (AI evaluation)

A standardised test or dataset used to measure and compare the performance of different AI models on a specific capability, such as maths problems, coding tasks, or general knowledge questions. Benchmarks are widely used in marketing and research alike, but are increasingly criticised for being "gamed" once a model's developers train on data resembling the test, or for measuring narrow skills that do not reflect real-world usefulness. (See also: Model card, Emergent capabilities)