AI GlossaryOTechnical words in the news
OSWorld-2.0
A benchmark that tests whether AI agents can independently complete multi-step tasks on a real computer screen
In plain words
OSWorld-2.0 is like an exam that hands an AI agent a real computer screen and checks whether it can click the mouse and press keys like a human to finish a task from start to finish. For example, it might have the agent open a browser to look up information, then copy that result into a spreadsheet and save it — a multi-step task carried out on an actual operating system screen, which is then graded on whether it was completed correctly all the way through.
While an ordinary test just asks you to pick the right answer, this kind of benchmark is much tougher: if the agent closes the wrong window or clicks the wrong menu partway through, it isn't simply a small deduction — it can cause the entire task to fail. Because it demands both the ability to visually understand a screen and the ability to actually operate it, it's used as a yardstick for whether an AI can go beyond just talking well and actually stand in for real work.
The "2.0" in the name simply means the test tasks and scoring methods have been refined further compared to the previous version. Scores on this benchmark are often released alongside the announcement of new AI models.
See also
Stories using this term
- Reddit developer releases 'Unswarm' to auto-switch between multiple local LLMsAI · 2026.08.23
- ComfyUI Open-Sources Local MCP ServerAI · 2026.08.22
- Hermes Agent Cuts 12 Browser Tools Down to OneAI · 2026.08.11
- Hermes Agent Declares "Fully Open to Forking and Self-Hosting"AI · 2026.08.21
- Hermes Agent gets a dedicated remote computer, priced at a few dollars a monthAI · 2026.08.21
- Apache Foundation incubates local-first AI agent tool 'Maka'AI · 2026.08.22
