AI GlossaryㅌTechnical words in the news
Terminal-Bench 4.0
A benchmark suite that scores how well AI models handle real-world tasks in a command-line computer environment
In plain words
Terminal-Bench 4.0 sits an AI model in front of a black-screen, text-only computer environment and grades how well it can find files, install programs, and fix problems. Think of it like handing a junior developer a real computer and saying 'go fix this bug,' then scoring them on the outcome.
This score matters because when companies release a new model, they don't just claim it's smarter than the last one — they back it up with this benchmark score. But the score isn't a pure measure of intelligence; it can also shift depending on how much a model's safety guardrails intervene.
How it shows up in the news
In articles, you might see something like: "Anthropic explained that the gap between the two models' Terminal-Bench 4.0 scores came from the older, less refined cyber safety guardrails interfering with some tasks." The easy mistake here is thinking the score gap means the models differ in raw 'skill' — when actually it means the tightness of the safety guardrails placed on the same model got in the way of task performance.
See also
Stories using this term
- DeepSeek V4 Pro GA Benchmarks Leak, Nears Top Open-Source TierAI · 2026.08.13
- DeepSeek-V4-Pro launches officially with three-tier reasoning intensity controlAI · 2026.08.14
- Shepherd, open-source runtime for rewinding agent execution unveiledAI · 2026.08.09
- Microsoft open-sources unit-testing agent that reads repositories and writes testsAI · 2026.08.09
- Perplexity's local agent beats Hermes, Pi in benchmarksAI · 2026.08.26
- GLM-5.3 API released, Terminal-Bench score jumps from 4.6 to 28.3AI · 2026.08.19
