METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryATechnical words in the news

ARC-AGI-3 benchmark

A test that measures an AI's reasoning ability by how well it solves puzzles it has never seen before

In plain words

The ARC-AGI-3 benchmark measures whether an AI can actually think, by having it solve puzzles it has never encountered before. It's a bit like a school exam that doesn't test how well you memorized the textbook, but how well you can solve a brand-new application problem on the spot. A higher score means the AI is better at spotting the underlying rules and solving problems in unfamiliar situations.

For example, if an AI model scored 13.3 on this test and later scored 38.3, that's read as a sign that it got better at adapting to entirely new types of problems, not just that it memorized more. That's why companies release this score alongside new models, using it as a way to show how much smarter the model has become.

How it shows up in the news

In the article it appears like this: "GPT-5.6 Sol reportedly raised its ARC-AGI-3 benchmark score from 13.3% to 38.3%, while cutting its output tokens to roughly a sixth of before." What's easy to misunderstand here is that this number doesn't measure a chatbot's eloquence or how much it knows — since the test measures how efficiently it solves puzzles it has never seen before, a rising score means it's solving harder problems with shorter answers.

See also

Stories using this term

Browse every entry