
Summary
- Upstage has released its in-house flagship reasoning model Solar Pro 4, which scored 42 on the Artificial Analysis Intelligence Index, a sharp jump from Solar Pro 3's 14 in April.
- On GDPval-AA v2, a benchmark of practical, real-world tasks, the model posted an Elo of 1276, surpassing the human baseline of 1000, and a gap of 778 points over its predecessor's 498.
- The knowledge-reliability metric AA-Omniscience improved from -53 to -1, but much of that gain came from the model attempting fewer answers rather than getting more of what it knows right.
A single number tells the story of this release. 498 to 1276. Upstage's new flagship model, Solar Pro 4, scored an Elo 778 points higher than its predecessor on a benchmark of practical, real-world tasks. That is an unusually large gap between models released by the same company just four months apart.

From 14 to 42
The evaluation was conducted by Artificial Analysis, an independent benchmarking firm. It combines results from multiple sub-tests into a single Intelligence Index, and Solar Pro 4 scored 42 — three times the 14 that Solar Pro 3 scored when it launched in April 2026. Solar Pro 4 is Upstage's new in-house (closed) flagship reasoning model, succeeding its predecessor.
For context: on the same index, xAI's Grok 4.6 scored 61, putting it in the same top tier as GPT-5.6 Sol. A score of 42 is not at the frontier's top rung, but the size of the jump a domestic lab made in a single generation stands out against the generational gains posted by frontier labs — Grok 4.5 to 4.6, for instance, gained just 5 points.

Real-world tasks: crossing the human baseline
The biggest improvement came in "agentic real-world work." GDPval-AA v2 gives models tasks drawn from actual jobs and compares their performance against a human baseline fixed at Elo 1000. Solar Pro 4 scored 1276, clearing that baseline. Artificial Analysis noted it comes in slightly ahead of MiMo-V2.5-Pro (1266). For the Qwen model cited alongside it, the post's body text lists Qwen3.7 Max at 1272, while the accompanying leaderboard table lists Qwen3.8 Max at 1737 — the discrepancy suggests the two references are to different versions.
| Model | GDPval-AA v2 Elo | |
|---|---|---|
| Claude Opus 5 (max) | 1849 | 100 |
| Grok 4.6 (high) | 1753 | 95 |
| GPT-5.6 Sol (max) | 1728 | 93 |
| Gemini 3.6 Flash | 1422 | 77 |
| Solar Pro 4 | 1276 | 69 |
| MiMo-V2.5-Pro | 1266 | 68 |
| Human baseline | 1000 | 54 |
| Solar Pro 3 | 498 | 27 |

Fewer hallucinations, but also fewer answers
The second metric needs some unpacking. AA-Omniscience measures a model's ability to distinguish what it knows from what it doesn't — rewarding correct answers, penalizing confident wrong answers (hallucinations), and not penalizing a refusal to answer. Solar Pro 4's score rose from -53 to -1.
However, Artificial Analysis pointed out that the gain came not from broader knowledge but from the model simply declining to answer more often. Where its predecessor attempted 92% of questions, Solar Pro 4 attempted only 41%. As a result, its hallucination rate fell from 88% to 24%. In effect, the model has learned to stay silent when it's likely to be wrong. That's a clear gain for reliability, but it's hard to read as evidence that the model's actual knowledge base has expanded.
| Metric | Solar Pro 3 | Solar Pro 4 |
|---|---|---|
| AA-Omniscience Index | -53 | -1 |
| Answer attempt rate | 92% | 41% |
| Hallucination rate | 88% | 24% |

Fewer words, but more time
Efficiency metrics were mixed. Output tokens used per task fell from 52k to 43k, a drop of about 17%. Even so, Artificial Analysis noted the model is still verbose compared to others at a similar intelligence level. More notable is time: despite using fewer tokens, time per task rose from 6.9 minutes to 9.5 minutes. Reasoning models churn through extended internal deliberation before producing an answer, and that process appears to have grown heavier.

What actually changes
A domestic lab's in-house model has now climbed above the human baseline on an independent evaluator's real-world task leaderboard. A score of 42 doesn't put it in the same league as Claude Opus 5 or the GPT-5.6 family, but the sheer size of the gap from its predecessor signals a major shift in reasoning training within a single generation. For practitioners, the more important change is one of character. A model with an 88% hallucination rate is hard to trust for document summarization or internal Q&A. One that has brought that down to 24% — by declining to answer rather than answering wrongly — at least reduces the cost of review. On the other hand, it may leave more than half of all questions unanswered, and response times have grown longer than before. Which tasks it's suited for should be decided with both of these traits in mind.





Comments