AI GlossaryㅌWords you meet while using AI
TerminalBench 3.0
A benchmark that tests how well an AI can solve coding problems on its own inside a terminal (command-line) environment.
In plain words
TerminalBench 3.0 is a test that measures how well an AI model can carry out real developer-style coding tasks inside a terminal — a screen where you can only type commands, with no buttons or menus. Given a problem, the AI has to type in commands on its own to solve it, and its result is graded as correct or not to produce a score.
Think of it like handing a chef only a storeroom of ingredients — no recipe, no appliances with buttons — and having them cook a dish entirely by hand, then judging how accurate the finished dish is. Since the AI can't click icons or menus and must direct every action through text alone, this benchmark gives a clean read on how autonomously it can solve problems without human help.
This score is often cited when comparing the abilities of AI coding tools. A higher score means the AI can carry a complex development task through to completion on its own, without human intervention.
How it shows up in the news
The article notes that Qwen3.8-Max-0902 scored 29.0 on TerminalBench 3.0, a big jump from its previous version's 11.3, but still short of Claude Opus 5's 42.7. A common misunderstanding here is that a lower score doesn't mean the AI's overall coding ability is weak. TerminalBench 3.0 is just one of several coding metrics — it specifically measures autonomy in completing tasks within a terminal environment. As the same article shows Qwen outperforming Claude Opus 5 on its own benchmark, QwenSWEBench V2, rankings can shift depending on which test is used.
See also
Stories using this term
- Qwen3.8-Max Gets Coding and Collaboration Boost With 0902 UpdateAI · 2026.09.02
- MiniMax Code updates browser automation and goal modeAI · 2026.08.12
- GLM-5.3 API released, Terminal-Bench score jumps from 4.6 to 28.3AI · 2026.08.19
- Shepherd, open-source runtime for rewinding agent execution unveiledAI · 2026.08.09
- Perplexity's local agent beats Hermes, Pi in benchmarksAI · 2026.08.26
- Meta's Muse Spark 1.3 beats GPT-5.6 and Opus 5 on coding benchmarksAI · 2026.09.03
