METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅌWords you meet while using AI

TerminalBench 3.0

A benchmark that tests how well an AI can solve coding problems on its own inside a terminal (command-line) environment.

In plain words

TerminalBench 3.0 is a test that measures how well an AI model can carry out real developer-style coding tasks inside a terminal — a screen where you can only type commands, with no buttons or menus. Given a problem, the AI has to type in commands on its own to solve it, and its result is graded as correct or not to produce a score.

Think of it like handing a chef only a storeroom of ingredients — no recipe, no appliances with buttons — and having them cook a dish entirely by hand, then judging how accurate the finished dish is. Since the AI can't click icons or menus and must direct every action through text alone, this benchmark gives a clean read on how autonomously it can solve problems without human help.

This score is often cited when comparing the abilities of AI coding tools. A higher score means the AI can carry a complex development task through to completion on its own, without human intervention.

How it shows up in the news

The article notes that Qwen3.8-Max-0902 scored 29.0 on TerminalBench 3.0, a big jump from its previous version's 11.3, but still short of Claude Opus 5's 42.7. A common misunderstanding here is that a lower score doesn't mean the AI's overall coding ability is weak. TerminalBench 3.0 is just one of several coding metrics — it specifically measures autonomy in completing tasks within a terminal environment. As the same article shows Qwen outperforming Claude Opus 5 on its own benchmark, QwenSWEBench V2, rankings can shift depending on which test is used.

See also

Stories using this term

Browse every entry