
이미지: X — 벤치마크·평가 화면 갈무리
Summary
- Upstage has released its self-developed flagship reasoning model, Solar Pro 4, which scored 42 on the Artificial Analysis Intelligence Index, a sharp jump from Solar Pro 3's score of 14 last April.
- On GDPval-AA v2, a benchmark for practical work tasks, it recorded an Elo of 1276, surpassing the human baseline of 1000 and marking a 778-point gain over its predecessor's score of 498.
- The knowledge-reliability metric AA-Omniscience improved from -53 to -1, but much of that gain came not from getting more known facts right, but from attempting fewer questions overall.
- 모델
- Solar Pro 4 (업스테이지, 자체 개발 비공개 플래그십 추론 모델)
- Artificial Analysis Intelligence Index
- 42점 (Solar Pro 3은 14점)
- 전작
- 2026년 4월 공개된 Solar Pro 3을 대체
- GDPval-AA v2 Elo
- 1276 (Solar Pro 3은 498, 인간 기준선 1000)
- AA-Omniscience Index
- -1 (Solar Pro 3은 -53)
- 응답 시도율·환각률
- 시도율 41%(전작 92%), 환각률 24%(전작 88%)
- 과제당 출력 토큰
- 43k (전작 52k, 약 17% 감소)
- 과제당 소요 시간
- 9.5분 (전작 6.9분)
A single number reveals the character of this release. 498 to 1276. Upstage's new flagship model, Solar Pro 4, posted an Elo score 778 points higher than its predecessor on a benchmark for practical, work-oriented tasks. That is an unusually large gap between models from the same company released just four months apart.

From 14 to 42
The evaluation was conducted by Artificial Analysis, an independent benchmarking organization. The firm combines multiple sub-tests into a single Intelligence Index score, and Solar Pro 4 scored 42 on it — triple the score of 14 that Solar Pro 3 received when it launched in April 2026. Solar Pro 4 is Upstage's new proprietary (closed) flagship reasoning model, replacing its predecessor.
For context, xAI's Grok 4.6 scored 61 on the same index, putting it in the same top tier as GPT-5.6 Sol. A score of 42 doesn't place Solar Pro 4 at the very top of the frontier, but the leap a domestic lab made in a single generation stands out against the generation-over-generation gains of frontier labs — Grok 4.5 to 4.6 improved by just 5 points.

A practical-task benchmark that beat the human baseline
The biggest improvement came in "agentic real-world work." GDPval-AA v2 assigns models tasks drawn from actual job functions and compares performance against a human baseline fixed at an Elo of 1000. Solar Pro 4 scored 1276, clearing that baseline. Artificial Analysis noted that this puts it slightly ahead of MiMo-V2.5-Pro (1266). Regarding the Qwen model referenced alongside it, the post text cites Qwen3.7 Max at 1272, while the attached leaderboard table lists Qwen3.8 Max at 1737 — suggesting the two references may point to different versions.

| Model | GDPval-AA v2 Elo | |
|---|---|---|
| Claude Opus 5 (max) | 1849 | 100 |
| Grok 4.6 (high) | 1753 | 95 |
| GPT-5.6 Sol (max) | 1728 | 93 |
| Gemini 3.6 Flash | 1422 | 77 |
| Solar Pro 4 | 1276 | 69 |
| MiMo-V2.5-Pro | 1266 | 68 |
| Human baseline | 1000 | 54 |
| Solar Pro 3 | 498 | 27 |

Hallucinations fell, but so did the willingness to answer
The second metric requires some interpretation. AA-Omniscience measures a model's ability to distinguish what it knows from what it doesn't. It rewards correct answers, penalizes confidently wrong (hallucinated) answers, and gives no penalty for declining to answer. Solar Pro 4 improved from -53 to -1 on this index.
However, Artificial Analysis pointed out that the gain stemmed not from broader knowledge but from the model answering less. While its predecessor attempted to answer 92% of questions, Solar Pro 4 attempted only 41%. As a result, its hallucination rate dropped from 88% to 24%. In effect, the model has been trained to stay silent rather than risk being wrong. That is a clear gain for reliability, but it's harder to read as evidence that the model's actual knowledge has expanded.
| Metric | Solar Pro 3 | Solar Pro 4 |
|---|---|---|
| AA-Omniscience Index | -53 | -1 |
| Answer attempt rate | 92% | 41% |
| Hallucination rate | 88% | 24% |

Fewer words, but more time
Efficiency metrics were mixed. Output tokens used per task dropped from 52k to 43k, a roughly 17% reduction. Even so, Artificial Analysis noted that Solar Pro 4 still uses more tokens than other models of comparable intelligence. More notable is the time factor: despite using fewer tokens, time spent per task rose from 6.9 minutes to 9.5 minutes. Since reasoning models run extended internal deliberation before producing an answer, this suggests that process has become heavier.

So what actually changes
A domestic lab's proprietary model has now cleared the human baseline on an independent evaluator's practical-work leaderboard. A score of 42 doesn't put it in the same league as Claude Opus 5 or the GPT-5.6 family, but the sheer size of the gap versus its predecessor signals a major shift in reasoning-training approach within a single generation. For practitioners, the more important shift is one of character. A model with an 88% hallucination rate is hard to trust for document summarization or internal Q&A. One that has brought that down to 24% — by declining to answer rather than answering wrongly — at least reduces the cost of review. The tradeoff is that it may fail to answer more than half of all questions, and response times have grown longer than its predecessor's. Deciding which tasks to assign it to requires weighing both of these traits together.



