Benchmark (AI evaluation)
Jump to navigation
Jump to search
Benchmark (AI evaluation)
A standardised test or dataset used to measure and compare the performance of different AI models on a specific capability, such as maths problems, coding tasks, or general knowledge questions. Benchmarks are widely used in marketing and research alike, but are increasingly criticised for being "gamed" once a model's developers train on data resembling the test, or for measuring narrow skills that do not reflect real-world usefulness. (See also: Model card, Emergent capabilities)