AI GlossaryㅇWords you meet while using AI
WorkArena Elo
An Elo-style ranking score, rated by human evaluators, that measures how well an AI model handles office work.
In plain words
WorkArena Elo is a ranking score that shows how well an AI model performs real office tasks, based on direct human judgment. Much like how an Elo rating boils down a chess player's skill into a single number, this score comes from having multiple AI models tackle 145 real-world tasks across 29 industries, with human evaluators comparing the results and assigning scores.
A higher number means better skill at office-style collaboration. What sets it apart from other benchmarks that measure coding ability is that it looks at how smoothly a model handles multi-day tasks — like drafting a press release, coordinating schedules, or organizing materials — rather than one-off, single-day jobs.
How it shows up in the news
Articles use it like this: "On the expert-rated collaboration index (WorkArena Elo), Qwen3.8-Max-0902 scored 1468, edging out Claude Opus 5 (1437) but falling short of GPT5.6 Sol (1482)." A common point of confusion is that this score does not represent coding ability. In the same article, coding benchmarks (like Terminal-Bench) and WorkArena Elo are listed as entirely separate categories — a model can lag in coding while still leading in office collaboration.
See also
Stories using this term
- Qwen3.8-Max Gets Coding and Collaboration Boost With 0902 UpdateAI · 2026.09.02
- NVIDIA releases Switchyard, an LLM routing proxy for coding agentsAI · 2026.08.20
- Meta Unveils First 10 Tasks in WildArtifactBench, a Benchmark for AI AgentsAI · 2026.08.21
- Reddit developer releases 'Unswarm' to auto-switch between multiple local LLMsAI · 2026.08.23
- Grok 4.6 unveiled: top-tier performance at half the priceAI · 2026.08.13
- Upstage Solar Pro 4 Intelligence Index jumps from 14 to 42AI · 2026.08.13
