What it is
An AI benchmark is a standardized test for models. It has a fixed set of tasks, a rule for scoring answers, and usually a leaderboard. Because every model faces the same questions, you can compare scores across models and across time.
Common kinds include multiple-choice knowledge exams, math word problems, coding challenges where the model must make failing tests pass, and agent tests where the model operates a computer to finish a job.
How it works
A benchmark is run by feeding the tasks to a model and checking its answers against a key or an automatic grader. The result is usually a percentage. Some benchmarks are graded by running code, some by comparing to a reference answer, and some by another model acting as judge.
Who runs the test matters a great deal:
- Self-reported: the lab runs the benchmark and publishes the score. Settings such as how many attempts the model gets, or how much it is allowed to think, are chosen by the lab.
- Independent: a third party runs the model under its own settings, which makes comparisons fairer but not perfect.
Why scores can mislead
A high score is useful evidence, and it is easy to over-read. A few common problems:
- Contamination: if test questions leaked into the model's training data, the score measures memory, not skill.
- Saturation: once most models score near the top, the test can no longer tell good from great.
- Narrow coverage: a benchmark measures its own tasks. A model that tops a coding test may still write poor emails.
- Settings: the same model can score very differently depending on its prompt, tools, thinking budget and number of tries.
How to use them
Treat a benchmark as a filter, not a verdict. Use it to shortlist models, then test the shortlist on your own work with a few real examples. Check whether the figure is self-reported or independent, and look for the settings behind it. When two models are close, price, speed and reliability on your tasks usually decide the choice better than a point of difference on a leaderboard.
An example
Suppose you need a model to classify support tickets. A reasoning benchmark may say little about that. Instead, collect fifty real tickets, label them yourself, and score each candidate model against your labels. That is a small private benchmark, and it often beats any public one.
Cost and speed belong in the comparison too. Pricing is counted in tokens, and long inputs depend on the model's context window. Safety testing such as red teaming is a different kind of evaluation, aimed at finding harmful behavior rather than scoring ability.