AI GlossaryATechnical words in the news
ARC-AGI-3 benchmark
A test that measures an AI's reasoning ability by how well it solves puzzles it has never seen before
In plain words
The ARC-AGI-3 benchmark measures whether an AI can actually think, by having it solve puzzles it has never encountered before. It's a bit like a school exam that doesn't test how well you memorized the textbook, but how well you can solve a brand-new application problem on the spot. A higher score means the AI is better at spotting the underlying rules and solving problems in unfamiliar situations.
For example, if an AI model scored 13.3 on this test and later scored 38.3, that's read as a sign that it got better at adapting to entirely new types of problems, not just that it memorized more. That's why companies release this score alongside new models, using it as a way to show how much smarter the model has become.
How it shows up in the news
In the article it appears like this: "GPT-5.6 Sol reportedly raised its ARC-AGI-3 benchmark score from 13.3% to 38.3%, while cutting its output tokens to roughly a sixth of before." What's easy to misunderstand here is that this number doesn't measure a chatbot's eloquence or how much it knows — since the test measures how efficiently it solves puzzles it has never seen before, a rising score means it's solving harder problems with shorter answers.
See also
Stories using this term
- GPT-6 Astra's first 48 hours bring real-world use from architecture to roboticsThe Lab · 2026.09.06
- Astra's 99.9% Score Came From the Harness, Not the ModelAI · 2026.09.04
- AI models talk faster without words, thanks to a Fields Medalist's bridgeThe Lab · 2026.09.03
- Prime Agent technical report shows ARC-AGI-3 score jump from 30% to 95.5%AI · 2026.08.27
- NVIDIA Swaps Harness, Lifts AI Agent Score from 30% to 100%AI · 2026.08.22
- OpenAI cuts GPT-5.6 Sol API pricing by over 20% for three monthsAI · 2026.08.22
