AI GlossaryㅂTechnical words in the news
Benchmark
A standardized test that AI models all take together. It's the score behind every "which model is smarter" headline.
In plain words
A benchmark is a standardized test that AI models all take under the same conditions. Just as students are compared by their exam scores, models are compared by having them solve the same set of problems. Every number cited in a "which model is smarter" article traces back to one of these tests.
The subjects vary widely. General knowledge (MMLU), PhD-level science (GPQA), competition math (AIME), real-world coding (SWE-bench) — those unfamiliar strings of letters attached to the bar charts in new model announcements are all test names.
Two things to watch for when reading these numbers. First, most published scores are "self-reported" — measured by the company under conditions favorable to itself, so they can differ from third-party verification. Second, being good at taking tests isn't the same as being good at doing the job. Disputes over "benchmark contamination," where leaked questions inflate scores without real improvement, come up often too. So it's more accurate to check which test was used and who measured it than to fixate on a few-point gap.
How it shows up in the news
"New model claims to beat rivals across all major benchmarks" — the word "claims" shows up because of the self-reporting problem above. METAL's benchmark section has explanations for each individual test.
Try it yourself
- Search for "LMArena" (Chatbot Arena) and go to the site — it's a blind comparison platform that shows answers from two AI models side by side without revealing their names.
- Ask any question you like and vote for the better answer. After voting, the models' identities are revealed.
- Millions of these votes accumulated together form the "Arena ranking" — a good way to see for yourself that exam scores (benchmarks) and how people actually feel about a model (Arena) can diverge.
See also
Stories using this term
- Liquid AI releases DSpark for its vision modelAI · 2026.09.25
- OpenAI publishes Harvey GPT-6 Astra case studyAI · 2026.09.24
- OpenAI Publishes V7 Context Graph Case StudyAI · 2026.09.22
- Specific Releases Real-SWE Enterprise Code BenchmarkAI · 2026.09.13
- Sakana AI Ships Two Models That Pick Other ModelsAI · 2026.09.12
- Silicon Valley forward-deployed engineer postings surge 1,000%Business · 2026.09.04
