A common test for comparing models
In AI, a benchmark is a standardized test for measuring a model’s capability. There are established problem sets for mathematics, programming, general knowledge, long-document comprehension, and more, and models are scored on the same questions. This is what launch announcements mean by a record score on a given benchmark.
What gets measured
- Knowledge and reasoning: university-level questions across fields, and hard logic or mathematics problems
- Coding: whether real software tasks can be completed
- Agentic ability: whether a job can be finished using tools
- Long-context comprehension, multilingual performance, and safety
As with human exams, every test has its own biases. The model ranked first overall is not necessarily the best one for your use.
Know the limits
Benchmarks attract teaching to the test. Contamination — test questions leaking into training data and inflating scores — is a documented problem, and high scores do not always match the experience of using a model. Tests also get saturated as models improve, so harder ones keep appearing in an endless cycle.
That is why evaluations closer to lived experience, such as ranking models by user votes, are used alongside them.
Reading scores sensibly
Benchmarks are an objective reference point, not an absolute measure. For choosing a model, running your own real task through it is still the most reliable test. When you see a score competition in the news, look at which area was tested and how large the gap actually is.
Results appear in vendor announcements and on leaderboard sites that collect models side by side. Looking at one occasionally gives you a sense of the current field.
You do not need to memorize benchmark names. Knowing that a standardized score competition exists is enough to read AI coverage.