METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅌTechnical words in the news

Terminal Bench 2.1

A benchmark that measures how well an AI agent can complete multi-step tasks on its own in a computer's command-line environment

In plain words

Terminal Bench 2.1 is a test that measures how well an AI agent actually gets work done in a computer's command-line screen (the terminal). Answering questions smoothly in a chat window and being dropped into an unfamiliar office to find, organize, and process documents in order are completely different skills — this test measures the latter.

Think of it like a cooking certification exam. It doesn't ask whether you know the names of ingredients; it grades whether you can actually stand in a kitchen and finish several dishes in the right order. The whole process — writing code, running commands, checking results, and deciding the next step on your own — is what gets evaluated.

A higher score means an automation tool built on that model is more likely to carry out complex tasks to completion with less human hand-holding. But this one test doesn't represent everything a model can do.

How it shows up in the news

It shows up in articles comparing rankings between models, as in "DeepSeek V4-Pro scored 87.9 on Terminal Bench 2.1, surpassing Opus-4.8 (85.0)." A common misunderstanding is treating this score as a measure of a model's overall intelligence — in reality, it's a narrow measure of practical skill at finishing tasks inside a terminal. Rankings can flip for the same model on other benchmarks (HLE, Cybergym, etc.).

Try it yourself

If you're using a tool with coding-agent features, try giving it a multi-step task instead of a single one-off question. For example, give it a prompt like "find the log files in this folder, filter out only the errors, and create compressed archives organized by date," and watch whether the agent works through it to the end on its own without getting stuck partway — that's a hands-on feel for what this test tries to measure.

See also

Stories using this term

Browse every entry