How are AI models actually evaluated, and what do benchmarks mean?
Against standardised test sets with published scores — and the honest position is that benchmarks measure something real, measure it narrowly, and are increasingly unreliable as headline comparisons for reasons that are structural rather than accidental.
The main approaches:
Static benchmarks — fixed question sets covering knowledge, reasoning, mathematics, coding and language understanding, scored automatically.
Human preference evaluation, where people compare outputs from two models blind and a rating emerges from many comparisons. Captures qualities automated tests miss, and reflects preference rather than correctness — which rewards confident, well-formatted answers.
Task-based evaluation, measuring completion of realistic multi-step work, which is closer to what people actually want to know.
Domain-specific evaluation by specialists in medicine, law or engineering.
Red teaming, deliberately attempting to elicit harmful or incorrect output.
Why headline scores mislead:
Contamination. If benchmark questions appear anywhere in training data, the model may have effectively seen the answers. This is extremely difficult to rule out at web scale, and is the single biggest problem with static benchmarks.
Saturation. Once a benchmark is widely used, models converge near the ceiling and it stops discriminating — and remaining errors may be errors in the benchmark itself.
Optimisation to the measure. When a benchmark becomes a target, it stops being a good measure, which is Goodhart's law applied to model development.
Narrow coverage. A score on multiple-choice questions says little about sustained reasoning, long-context work, tool use or reliability.
Missing what matters operationally — latency, cost, refusal behaviour, consistency across runs, and failure modes.
What actually tells you whether a model suits you:
Build your own evaluation set from real examples of your task, with known good answers. This is the most valuable thing available, and almost nobody does it.
Test the failure cases you care about, not the average.
Measure cost and latency alongside quality.
Re-evaluate after model updates, since behaviour changes.