
Summary
- xAI's Grok 4.6 (listed as "SpaceXAI's Grok 4.6") scored 61 on the Artificial Analysis Intelligence Index, putting it in the same top tier as GPT-5.6 Sol
- Across three practical benchmarks — GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1 — it matched or beat every model except Claude Opus 5
- Pricing stayed the same as the previous Grok 4.5, and per-task cost is on par with Kimi K3, far below Claude Opus 5 and GPT-5.6 Sol
Five points in a month — Grok's rapid comeback
xAI's large language model Grok 4.6 scored 61 on the Intelligence Index from evaluator Artificial Analysis. The post in question labeled the model "SpaceXAI's Grok 4.6." A score of 61 puts it on par with GPT-5.6 Sol, placing it among the top-tier models. What stands out is the speed: the score jumped 5 points in just over a month since the previous Grok 4.5 was released.

Booking the runner-up spot behind Claude Opus 5
Grok 4.6 showed strength in practical, work-oriented tasks such as knowledge work, terminal operation, and customer service. Across three benchmarks measured independently by Artificial Analysis — GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1 — Grok 4.6 matched or outperformed every model except Claude Opus 5.
| Model | GDPval-AA v2 (Elo) | τ³-Banking | Terminal-Bench v2.1 |
|---|---|---|---|
| Claude Opus 5 (max) | 1849 | 42.1% | 89.1% |
| Grok 4.6 (high) | 1753 | 50.7% | 88.4% |
| Claude Fable 5 | 1741 | 38.1% | 84.6% |
| Qwen3.8 Max | 1737 | 51.3% | 81.3% |
| GPT-5.6 Sol (max) | 1728 | 44.3% | 88.0% |
| Kimi K3 (max) | 1682 | 46.0% | 85.0% |
GDPval-AA v2 grades real-world work tasks against a human baseline (1000 points). τ³-Banking measures banking-scenario tasks, and Terminal-Bench v2.1 measures the ability to operate a computer terminal. Grok 4.6 held the No. 2 spot across all three metrics.

Cheap and capable — the cost-performance angle
Grok 4.6's standard pricing is unchanged from Grok 4.5. Yet its intelligence score rose 5 points. According to Artificial Analysis, its per-task cost is comparable to Kimi K3 and far lower than Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5. Artificial Analysis assessed that when performance and cost are plotted together, the model lands on the Pareto frontier.
The model also made its first appearance on AA-Briefcase, a private benchmark covering long-duration knowledge work tasks. It scored an Elo of 1577 there, again ranking second behind Claude Opus 5. Artificial Analysis noted that it "consistently performed strongly across grading criteria, document structure, and analysis quality."

What the Intelligence Index means
The Intelligence Index combines scores from multiple benchmarks into a single ranking of models. Meta's Muse Spark 1.2, released on August 8, previously scored 54 on this index, tying for third place with SpaceXAI. By contrast, Ant Group's open-weight small model Ling 3.0 Tiny, released August 12, scored only 25. Seen in this context, Grok 4.6's score of 61 places it in the crowded top tier dominated by large closed models, while the gap with open-weight small models remains wide.

So what changes
Until now, top-tier intelligence and low cost were seen as difficult to achieve together. Grok 4.6 kept its price unchanged while raising performance, delivering both at once. For enterprises that need to run long agentic workloads repeatedly, this adds another option that offers top-tier performance without the cost burden — even if it falls short of Claude Opus 5's absolute best quality. It can be read as a sign that the axis of frontier model competition is shifting from "who is smartest" to "who can do more at the same cost."





Comments