AI GlossarySTechnical words in the news
SWE-bench Verified
A benchmark that automatically scores how well an AI coding assistant fixes real bugs from actual open-source projects, using human-verified problems.
In plain words
SWE-bench Verified is a set of problems that tests how well an AI coding assistant can fix software bugs that actually happened in the real world. Each problem comes with a real bug report, the code that needs fixing, and an automated test that checks whether the fix works — so the AI's answer is graded by whether it passes the test, not by a human reading through it.
The "Verified" in the name means that humans went back through the original, larger set of problems and manually checked each one, filtering out any that weren't actually solvable or that had unclear descriptions. What's left is this curated subset. Because of that extra scrutiny, scores on this benchmark are often cited as evidence for how much you can trust an AI to handle real bug-fixing work in a company's codebase.
How it shows up in the news
Articles use it as a reference point for showing off a new agent framework or model's capabilities, for example: "NVIDIA reported that NOOA achieved higher accuracy and lower token costs than existing harnesses across three benchmarks — SWE-bench Verified, CyberGym L1, and ARC-AGI-3."
A common misunderstanding: a high score here doesn't mean the model is good at all coding tasks. The problems all come from a specific set of open-source projects, so performance outside that scope can differ.
Try it yourself
When comparing coding assistants, questions like these can help:
What's this model's SWE-bench Verified score, and how many problems out of the total did it solve? How does that score compare to other coding assistants?
Don't just look at the score itself — checking how many problems it was measured against also helps you gauge how reliable that number really is.
See also
Stories using this term
- Open-source Ornith-1.5 claims scores on par with Claude OpusAI · 2026.08.21
- Qwen3.8-27B released as open weights under Apache 2.0AI · 2026.08.15
- Microsoft's new coding model falls short of DeepSeek on both price and performanceAI · 2026.08.12
- NVIDIA Unveils NOOA, a Framework That Builds Agents as Python ClassesAI · 2026.08.09
- Ai2 finds BBQ safety benchmark actually measures reasoning abilityAI · 2026.09.02
- Meta Unveils First 10 Tasks in WildArtifactBench, a Benchmark for AI AgentsAI · 2026.08.21
