METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅇTechnical words in the news

Elo rating

A ranking score system that measures relative skill by pitting two subjects head-to-head and adjusting scores based on who wins or loses

In plain words

Elo rating was originally devised to rank chess players' skill levels. Rather than grading an absolute score, it pits two players against each other and adjusts rankings based solely on who won and who lost. Beat a strong opponent, and your score jumps a lot; lose to a weak one, and it drops sharply.

The AI industry has borrowed this same method to evaluate outputs. For example, two AI agents each perform the same task, and a judge simply picks which of the two results is better, without scoring either one on an absolute scale. Repeat this comparison hundreds of times, and you can rank even tasks with no single correct answer—like video editing or 3D modeling, where it's hard to say outright what's "right" or "wrong."

Unlike tests with a fixed answer key, Elo rating is a relative measure. It doesn't tell you "how good is this agent on some absolute scale" but rather "is this agent better than that other agent." That makes it especially useful for evaluating complex, real-world tasks that don't have a predetermined correct answer.

How it shows up in the news

WildArtifactBench, an evaluation framework released by Meta AI, pits agents' outputs against each other head-to-head to calculate win rates and Elo ratings. Contrary to a common misconception, it doesn't grade outputs against a fixed correct answer like a test—it determines which agent's output is relatively better by comparing it directly against another agent's output.

See also

Stories using this term

Browse every entry