METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AA-AnalystAgent Benchmark: Claude Opus 5 Leads at 54%

Artificial Analysis tests AI on real-world spreadsheets and documents. Claude Opus 5 scored 54%; 57% of failures anchored on a wrong early hypothesis.

AA-AnalystAgent Benchmark: Claude Opus 5 Leads at 54%

Summary

  • Artificial Analysis has released AA-AnalystAgent, a benchmark for quantitative analysis agents working with real spreadsheets and documents, with Claude Opus 5 ranking first at 54%, GPT-5.5 at 50%, and Claude Fable 5 at 49%.
  • On single-attempt accuracy (pass@1), GPT-5.5 (xhigh) led with 66%, but on pass^5, which requires success across all five attempts, Claude Opus 5 came out ahead.
  • An analysis of 1,567 failed attempts found that 57% involved "fixation on an incorrect initial hypothesis," while models scoring the same 20% differed in cost per task by as much as $0.05 versus $1.34.

The same analytical task was run five times. If a model got four right and one wrong, that problem scored zero. That's the grading method behind AA-AnalystAgent, a new benchmark from independent evaluator Artificial Analysis. Even at the top of the leaderboard, scores barely cleared the halfway mark. Claude Opus 5 took first place with 54%, followed by GPT-5.5 at 50% and Claude Fable 5 at 49%.

여러 AI 모델별로 첫 시도, 다섯 번 시도, 다섯 번 모두 통과 비율을 비교한 세로 막대그래프
이미지: @ArtificialAnlys (X)

A test that uses real files, not a question bank

Language model evaluations to date have mostly been written exams with fixed questions and answers in text form. AA-AnalystAgent is different. It hands models real-world spreadsheets and documents, requiring them to use tools to open files, extract numbers, and complete calculations. This kind of "test that requires using tools" is known as an agentic benchmark.

Artificial Analysis explained its reasoning for building this task by pointing to the nature of analyst work itself. "In real analyst work, professional judgment and expertise matter just as much," the team noted. In other words, getting a single correct answer is a different skill from producing analysis that's ready to be handed off to someone else.

ModelScore
Claude Opus 554%54
GPT-5.550%50
Claude Fable 549%49
Kimi K3 (max, top open-weight model)39%39
AI 모델별 실패 모드별 실패 시도 비율을 색상으로 나타낸 열 지도 그래프
이미지: @ArtificialAnlys (X)

What separated the ranks wasn't raw performance, but consistency

One striking detail is that single-attempt accuracy and final ranking don't line up. On pass@1 (single-attempt accuracy), GPT-5.5 (xhigh) led with 66%, closely followed by Gemini 3.1 Pro Preview and Claude Opus 5 (max) at 64%. But on pass^5 — which the evaluation team calls "pass-all-5," requiring success on all five attempts — Opus 5 pulled ahead. It wasn't the best-performing model that rose to the top, but the most consistent one.

MetricLeaderScore
pass@1 (single-attempt success rate)GPT-5.5 (xhigh)66%
pass@1 tied for 2ndGemini 3.1 Pro Preview / Claude Opus 5 (max)64% each
pass^5 (success on all 5 attempts)Claude Opus 5Leaderboard top

The announcement post did not specify exactly how the headline 54% score was calculated. But the divergence between pass@1 and the final score itself reveals something about what this leaderboard is actually measuring.

AI 모델별 실패 모드별 중간값 대비 실패 비율 차이를 색상으로 나타낸 열 지도 그래프
이미지: @ArtificialAnlys (X)

57% of failures: "couldn't let go of the initial hypothesis"

The evaluation team collected 1,567 failed attempts across 10 leading models and sorted them into seven failure categories, allowing multiple categories to apply to a single attempt. The most common failure mode — persisting with an incorrect hypothesis formed early on — appeared in 57% of all failures.

It's a familiar scene for anyone who has worked in the field. A model misreads the data, commits to a direction, and then forces every subsequent calculation to fit that direction. New human analysts do the same thing. The difference is that a person tends to pause and think "something's off," whereas a model plows ahead to the end with seemingly sound logic.

AI 모델별 다섯 번 모두 통과 비율과 작업당 비용을 비교한 산점도 그래프
이미지: @ArtificialAnlys (X)

Same score, 27x difference in cost

Layering in the cost dimension changes the picture again. Claude Sonnet 4.6 (max) and MiMo-V2.5-Pro both scored 20%, yet cost per task was $1.34 and $0.05 respectively — a 27x gap for the same result. The pattern holds at the top too: the most expensive model tested, Claude Opus 4.7 (max), cost $1.98 per task.

Among open-weight models, Kimi K3 (max) topped the field at launch with 39%, trailing the leading closed frontier models by 15 points while outperforming all eight other open-weight models tested alongside it. DeepSeek's DeepSeek V4 Flash came next.

AI 모델별 오픈 가중치와 독점 가중치 모델의 다섯 번 모두 통과 비율을 비교한 세로 막대그래프
이미지: @ArtificialAnlys (X)

So what does this change

On August 10, Artificial Analysis opened early access sign-ups for what it described as a developer-focused benchmarking product suite it's building. AA-AnalystAgent sits within that effort — an evaluation that turns office work itself directly into a test. For organizations looking to deploy AI in roles like finance, accounting, or research, where spreadsheets are the core of the job, there's now one more data point to consult.

The takeaway is simple: today's agents can pull off a convincing performance once, but not five times in a row. It's still too early to take their output at face value without review. The agent harness used to run the tasks — Stirrup, the execution shell that lets a model open files and run code — has been open-sourced on GitHub, making it possible to run the same tests independently rather than simply trusting someone else's reported results.

Comments