AI GlossaryMTechnical words in the news
MMLU-Pro
A benchmark exam that tests how well an AI model handles broad knowledge and reasoning across fields, using harder questions.
In plain words
MMLU-Pro is one of the exams used to measure an AI model's abilities. It mixes questions from many subjects—math, law, medicine, engineering—and scores models on how many they get right. Just as people prove their skills by passing certification exams, AI companies point to scores like this when launching a new model to claim "our model is this smart."
The "Pro" in the name means it's made harder than the original test. The earlier exam only had 4 answer choices, so guessing could sometimes get you the right answer. This version increases the choices to 10 and makes the questions themselves trickier, so you actually have to understand the material to solve them. That's why it's used to address the earlier exam's limitation, where scores between models often failed to show much difference.
Still, an exam score doesn't tell you everything about real ability. Some models may have memorized exam questions that leaked into their training data, or may be strong on tests but weak at real-world tasks. So it's safer to look at results from several benchmarks together rather than judging by a single score.
See also
Stories using this term
- Ai2 finds BBQ safety benchmark actually measures reasoning abilityAI · 2026.09.02
- Open-Weight AI Becomes Cheapest Shield and Easiest Spear in Same MonthThe Lab · 2026.09.02
- Multiverse Computing shrinks a model to 4-bit and gets one smarter than the originalAI · 2026.08.25
- Even the best AI models can't score above 60% on pure visionAI · 2026.08.15
- Hume AI Measures Benchmark Memorization in Speech Recognition ModelsAI · 2026.08.22
- Liquid AI's 300M Draft Model Speeds Up Decoding by Up to 3.18xAI · 2026.08.21
