METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossarySTechnical words in the news

SWE-bench Verified

A benchmark that automatically scores how well an AI coding assistant fixes real bugs from actual open-source projects, using human-verified problems.

In plain words

SWE-bench Verified is a set of problems that tests how well an AI coding assistant can fix software bugs that actually happened in the real world. Each problem comes with a real bug report, the code that needs fixing, and an automated test that checks whether the fix works — so the AI's answer is graded by whether it passes the test, not by a human reading through it.

The "Verified" in the name means that humans went back through the original, larger set of problems and manually checked each one, filtering out any that weren't actually solvable or that had unclear descriptions. What's left is this curated subset. Because of that extra scrutiny, scores on this benchmark are often cited as evidence for how much you can trust an AI to handle real bug-fixing work in a company's codebase.

How it shows up in the news

Articles use it as a reference point for showing off a new agent framework or model's capabilities, for example: "NVIDIA reported that NOOA achieved higher accuracy and lower token costs than existing harnesses across three benchmarks — SWE-bench Verified, CyberGym L1, and ARC-AGI-3."

A common misunderstanding: a high score here doesn't mean the model is good at all coding tasks. The problems all come from a specific set of open-source projects, so performance outside that scope can differ.

Try it yourself

When comparing coding assistants, questions like these can help:

What's this model's SWE-bench Verified score, and how many problems out of the total did it solve? How does that score compare to other coding assistants?

Don't just look at the score itself — checking how many problems it was measured against also helps you gauge how reliable that number really is.

See also

Stories using this term

Browse every entry