
Summary
- Artificial Analysis has released AA-AnalystAgent, a benchmark for quantitative analysis agents working with real spreadsheets and documents, with Claude Opus 5 ranking first at 54%, GPT-5.5 at 50%, and Claude Fable 5 at 49%.
- On single-attempt accuracy (pass@1), GPT-5.5 (xhigh) led with 66%, but on pass^5, which requires success across all five attempts, Claude Opus 5 came out ahead.
- An analysis of 1,567 failed attempts found that 57% involved "fixation on an incorrect initial hypothesis," while models scoring the same 20% differed in cost per task by as much as $0.05 versus $1.34.
The same analytical task was run five times. If a model got four right and one wrong, that problem scored zero. That's the grading method behind AA-AnalystAgent, a new benchmark from independent evaluator Artificial Analysis. Even at the top of the leaderboard, scores barely cleared the halfway mark. Claude Opus 5 took first place with 54%, followed by GPT-5.5 at 50% and Claude Fable 5 at 49%.

A test that uses real files, not a question bank
Language model evaluations to date have mostly been written exams with fixed questions and answers in text form. AA-AnalystAgent is different. It hands models real-world spreadsheets and documents, requiring them to use tools to open files, extract numbers, and complete calculations. This kind of "test that requires using tools" is known as an agentic benchmark.
Artificial Analysis explained its reasoning for building this task by pointing to the nature of analyst work itself. "In real analyst work, professional judgment and expertise matter just as much," the team noted. In other words, getting a single correct answer is a different skill from producing analysis that's ready to be handed off to someone else.
| Model | Score | |
|---|---|---|
| Claude Opus 5 | 54% | |
| GPT-5.5 | 50% | |
| Claude Fable 5 | 49% | |
| Kimi K3 (max, top open-weight model) | 39% |

What separated the ranks wasn't raw performance, but consistency
One striking detail is that single-attempt accuracy and final ranking don't line up. On pass@1 (single-attempt accuracy), GPT-5.5 (xhigh) led with 66%, closely followed by Gemini 3.1 Pro Preview and Claude Opus 5 (max) at 64%. But on pass^5 — which the evaluation team calls "pass-all-5," requiring success on all five attempts — Opus 5 pulled ahead. It wasn't the best-performing model that rose to the top, but the most consistent one.
| Metric | Leader | Score |
|---|---|---|
| pass@1 (single-attempt success rate) | GPT-5.5 (xhigh) | 66% |
| pass@1 tied for 2nd | Gemini 3.1 Pro Preview / Claude Opus 5 (max) | 64% each |
| pass^5 (success on all 5 attempts) | Claude Opus 5 | Leaderboard top |
The announcement post did not specify exactly how the headline 54% score was calculated. But the divergence between pass@1 and the final score itself reveals something about what this leaderboard is actually measuring.

57% of failures: "couldn't let go of the initial hypothesis"
The evaluation team collected 1,567 failed attempts across 10 leading models and sorted them into seven failure categories, allowing multiple categories to apply to a single attempt. The most common failure mode — persisting with an incorrect hypothesis formed early on — appeared in 57% of all failures.
It's a familiar scene for anyone who has worked in the field. A model misreads the data, commits to a direction, and then forces every subsequent calculation to fit that direction. New human analysts do the same thing. The difference is that a person tends to pause and think "something's off," whereas a model plows ahead to the end with seemingly sound logic.

Same score, 27x difference in cost
Layering in the cost dimension changes the picture again. Claude Sonnet 4.6 (max) and MiMo-V2.5-Pro both scored 20%, yet cost per task was $1.34 and $0.05 respectively — a 27x gap for the same result. The pattern holds at the top too: the most expensive model tested, Claude Opus 4.7 (max), cost $1.98 per task.
Among open-weight models, Kimi K3 (max) topped the field at launch with 39%, trailing the leading closed frontier models by 15 points while outperforming all eight other open-weight models tested alongside it. DeepSeek's DeepSeek V4 Flash came next.

So what does this change
On August 10, Artificial Analysis opened early access sign-ups for what it described as a developer-focused benchmarking product suite it's building. AA-AnalystAgent sits within that effort — an evaluation that turns office work itself directly into a test. For organizations looking to deploy AI in roles like finance, accounting, or research, where spreadsheets are the core of the job, there's now one more data point to consult.
The takeaway is simple: today's agents can pull off a convincing performance once, but not five times in a row. It's still too early to take their output at face value without review. The agent harness used to run the tasks — Stirrup, the execution shell that lets a model open files and run code — has been open-sourced on GitHub, making it possible to run the same tests independently rather than simply trusting someone else's reported results.





Comments