AI GlossaryGTechnical words in the news
GDPval-AA v2
A practical benchmark that tests AI on real job tasks and compares its performance against a human baseline score (Elo 1000)
In plain words
GDPval-AA v2 is a test that gives AI real workplace tasks and compares the results to scores from humans doing the same work. The level of human performance is fixed at 1000 points, and if an AI scores higher than that, it has cleared this baseline.
Think of it like a job performance review for a new hire. Tasks such as writing reports or organizing materials are assigned, and the results from an experienced worker are set as the reference point (1000 points). The AI's output is then placed side by side and scored against that reference. A score above 1000 means that, at least on this test, the AI outperformed the human average.
The evaluation is run by the independent organization Artificial Analysis, and the latest models from various AI companies are put through this test and ranked. That said, doing well on a single test doesn't mean an AI can fully replace real-world work.
How it shows up in the news
The article reports that "Upstage's Solar Pro 4 scored 1276 on GDPval-AA v2, rising above the human baseline (1000)." A common misunderstanding here is that this score doesn't fully prove the AI is 'better at the job than humans' across the board. It's more accurate to read it as meaning the AI exceeded the scoring criteria for the specific types of tasks included in the test.
See also
Stories using this term
- Meta's Muse Spark 1.3 beats GPT-5.6 and Opus 5 on coding benchmarksAI · 2026.09.03
- Google unveils Gemini 3.8 Flash, tuned for coding and agentic workAI · 2026.09.03
- Anthropic launches Fable 5.1 and Mythos 5.1AI · 2026.09.02
- Ox Alpha turns out to be GLM-5.3-FlashAI · 2026.08.26
- Upstage Solar Pro 4 Intelligence Index jumps from 14 to 42AI · 2026.08.13
- Musk Teases Grok 4.7 Will "Surpass Every Existing Model"AI · 2026.08.13
