METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Upstage Solar Pro 4 Intelligence Index jumps from 14 to 42

Score triples over predecessor in Artificial Analysis evaluation, real-world task Elo rises from 498 to 1276, surpassing human baseline

Upstage Solar Pro 4 Intelligence Index jumps from 14 to 42

Summary

  • Upstage has released its in-house flagship reasoning model Solar Pro 4, which scored 42 on the Artificial Analysis Intelligence Index, a sharp jump from Solar Pro 3's 14 in April.
  • On GDPval-AA v2, a benchmark of practical, real-world tasks, the model posted an Elo of 1276, surpassing the human baseline of 1000, and a gap of 778 points over its predecessor's 498.
  • The knowledge-reliability metric AA-Omniscience improved from -53 to -1, but much of that gain came from the model attempting fewer answers rather than getting more of what it knows right.

A single number tells the story of this release. 498 to 1276. Upstage's new flagship model, Solar Pro 4, scored an Elo 778 points higher than its predecessor on a benchmark of practical, real-world tasks. That is an unusually large gap between models released by the same company just four months apart.

여러 AI 모델의 GDPval-AA v2 점수를 막대그래프로 비교한 리더보드 차트
이미지: @ArtificialAnlys (X)

From 14 to 42

The evaluation was conducted by Artificial Analysis, an independent benchmarking firm. It combines results from multiple sub-tests into a single Intelligence Index, and Solar Pro 4 scored 42 — three times the 14 that Solar Pro 3 scored when it launched in April 2026. Solar Pro 4 is Upstage's new in-house (closed) flagship reasoning model, succeeding its predecessor.

For context: on the same index, xAI's Grok 4.6 scored 61, putting it in the same top tier as GPT-5.6 Sol. A score of 42 is not at the frontier's top rung, but the size of the jump a domestic lab made in a single generation stands out against the generational gains posted by frontier labs — Grok 4.5 to 4.6, for instance, gained just 5 points.

AA-Omniscience 지수, 정확도, 환각률을 각각 막대그래프로 나타낸 세 부분 차트
이미지: @ArtificialAnlys (X)

Real-world tasks: crossing the human baseline

The biggest improvement came in "agentic real-world work." GDPval-AA v2 gives models tasks drawn from actual jobs and compares their performance against a human baseline fixed at Elo 1000. Solar Pro 4 scored 1276, clearing that baseline. Artificial Analysis noted it comes in slightly ahead of MiMo-V2.5-Pro (1266). For the Qwen model cited alongside it, the post's body text lists Qwen3.7 Max at 1272, while the accompanying leaderboard table lists Qwen3.8 Max at 1737 — the discrepancy suggests the two references are to different versions.

ModelGDPval-AA v2 Elo
Claude Opus 5 (max)1849100
Grok 4.6 (high)175395
GPT-5.6 Sol (max)172893
Gemini 3.6 Flash142277
Solar Pro 4127669
MiMo-V2.5-Pro126668
Human baseline100054
Solar Pro 349827
AI 모델별 지능지수 작업당 출력 토큰 수를 답변과 추론으로 구분해 표시한 막대그래프
이미지: @ArtificialAnlys (X)

Fewer hallucinations, but also fewer answers

The second metric needs some unpacking. AA-Omniscience measures a model's ability to distinguish what it knows from what it doesn't — rewarding correct answers, penalizing confident wrong answers (hallucinations), and not penalizing a refusal to answer. Solar Pro 4's score rose from -53 to -1.

However, Artificial Analysis pointed out that the gain came not from broader knowledge but from the model simply declining to answer more often. Where its predecessor attempted 92% of questions, Solar Pro 4 attempted only 41%. As a result, its hallucination rate fell from 88% to 24%. In effect, the model has learned to stay silent when it's likely to be wrong. That's a clear gain for reliability, but it's hard to read as evidence that the model's actual knowledge base has expanded.

MetricSolar Pro 3Solar Pro 4
AA-Omniscience Index-53-1
Answer attempt rate92%41%
Hallucination rate88%24%
AI 모델별 지능지수 작업당 소요 시간과 지능지수 간 관계를 산점도로 나타낸 그래프
이미지: @ArtificialAnlys (X)

Fewer words, but more time

Efficiency metrics were mixed. Output tokens used per task fell from 52k to 43k, a drop of about 17%. Even so, Artificial Analysis noted the model is still verbose compared to others at a similar intelligence level. More notable is time: despite using fewer tokens, time per task rose from 6.9 minutes to 9.5 minutes. Reasoning models churn through extended internal deliberation before producing an answer, and that process appears to have grown heavier.

여러 AI 모델의 다양한 지능 평가 항목별 점수를 막대그래프로 비교한 종합 평가 차트
이미지: @ArtificialAnlys (X)

What actually changes

A domestic lab's in-house model has now climbed above the human baseline on an independent evaluator's real-world task leaderboard. A score of 42 doesn't put it in the same league as Claude Opus 5 or the GPT-5.6 family, but the sheer size of the gap from its predecessor signals a major shift in reasoning training within a single generation. For practitioners, the more important change is one of character. A model with an 88% hallucination rate is hard to trust for document summarization or internal Q&A. One that has brought that down to 24% — by declining to answer rather than answering wrongly — at least reduces the cost of review. On the other hand, it may leave more than half of all questions unanswered, and response times have grown longer than before. Which tasks it's suited for should be decided with both of these traits in mind.

Comments