Astra's 99.9% Score Came From the Harness, Not the Model
OpenAI's ARC-AGI-3 results sheet lists 62.7% and 99.9% side by side for the same model. Tracing why one model ends up with two scores reveals who actually deserves credit for this result.
350
Runs ARC-AGI, an abstract reasoning benchmark.
Current rank (1M)
66
Rank over the last 6 days
2026.09.08 – 2026.09.13 · High 65 · Low 74 · Now 66
OpenAI's ARC-AGI-3 results sheet lists 62.7% and 99.9% side by side for the same model. Tracing why one model ends up with two scores reveals who actually deserves credit for this result.
When we connect two models, we've always had one write text for the other to read.
Prime Intellect says the gain came from changing the execution shell around the model, not the model itself. The code is on GitHub.
That's the last story.