
이미지: X — 벤치마크·평가 화면 갈무리
Summary
- xAI's Grok 4.6 (listed as "SpaceXAI's Grok 4.6") scored 61 on the Artificial Analysis Intelligence Index, putting it in the top tier alongside GPT-5.6 Sol
- Across three practical benchmarks — GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1 — it matched or outperformed every model except Claude Opus 5
- Pricing stayed the same as its predecessor Grok 4.5, with per-task cost on par with Kimi K3 and well below Claude Opus 5 and GPT-5.6 Sol
- 모델명
- Grok 4.6 (보도에서는 SpaceXAI 소속으로 표기)
- Intelligence Index 점수
- 61점, GPT-5.6 Sol과 동급
- 전작 대비
- Grok 4.5 대비 +5점, 출시 약 한 달 만
- GDPval-AA v2 Elo
- 1753 (Claude Opus 5 1849에 이어 2위)
- τ³-Banking 점수
- 50.7% (Qwen3.8 Max 51.3%에 이어 2위)
- Terminal-Bench v2.1
- 88.4% (Claude Opus 5 89.1%에 이어 2위)
- AA-Briefcase Elo
- 1577, Claude Opus 5에 이어 2위
- 가격
- Grok 4.5와 동일, 과제당 비용은 Kimi K3와 비슷한 수준
Five points in a month — Grok's rapid comeback
xAI's large language model Grok 4.6 scored 61 on the Intelligence Index run by evaluator Artificial Analysis. The post in question labeled the model "SpaceXAI's Grok 4.6." A score of 61 puts it on par with GPT-5.6 Sol, placing it among the top-tier models. What stands out is the speed: the score rose 5 points in roughly a month since its predecessor, Grok 4.5, was released.

Second place, right behind Claude Opus 5
Grok 4.6 showed particular strength on practical, work-oriented tasks such as knowledge work, terminal operation, and customer service. Across three benchmarks measured independently by Artificial Analysis — GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1 — Grok 4.6 matched or outperformed every model except Claude Opus 5.
| Model | GDPval-AA v2 (Elo) | τ³-Banking | Terminal-Bench v2.1 |
|---|---|---|---|
| Claude Opus 5 (max) | 1849 | 42.1% | 89.1% |
| Grok 4.6 (high) | 1753 | 50.7% | 88.4% |
| Claude Fable 5 | 1741 | 38.1% | 84.6% |
| Qwen3.8 Max | 1737 | 51.3% | 81.3% |
| GPT-5.6 Sol (max) | 1728 | 44.3% | 88.0% |
| Kimi K3 (max) | 1682 | 46.0% | 85.0% |
GDPval-AA v2 grades real-world tasks against a human baseline (1000 points). τ³-Banking measures banking-scenario performance, and Terminal-Bench v2.1 measures the ability to operate a computer terminal. Grok 4.6 held second place across all three metrics.

Cheap and capable — cost-performance
Grok 4.6's standard pricing is unchanged from Grok 4.5. Yet its intelligence score rose by 5 points. According to Artificial Analysis, per-task cost is comparable to Kimi K3 and much lower than Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5. Artificial Analysis assessed that when performance and cost are plotted together, the model lands on the Pareto frontier.
The model also made its first appearance on AA-Briefcase, a private benchmark for long-horizon knowledge-work tasks. There it scored an Elo of 1577, again second only to Claude Opus 5. Artificial Analysis noted that it "showed consistently strong performance across grading criteria, document structure, and analytical quality."

What the Intelligence Index means
The Intelligence Index combines scores from multiple benchmarks into a single ranking across models. Meta's Muse Spark 1.2, released on August 8, scored 54 on this index, tying with SpaceXAI for third place at the time. By contrast, Ant Group's open-weight small model Ling 3.0 Tiny, released August 12, scored only 25. In that context, Grok 4.6's score of 61 signals entry into the crowded top tier dominated by large closed models, while the gap with open-weight small models remains wide.

What actually changes
Top-tier intelligence and low cost have long been seen as difficult to combine. Grok 4.6 achieves both by holding price steady while raising performance. For enterprises that need to run long-horizon agentic tasks repeatedly, this adds an option that delivers near-top-tier performance without the cost burden — even if it falls short of Claude Opus 5's absolute best quality. It can be read as a sign that competition among frontier models is shifting from "who's smartest" to "who accomplishes more at the same cost."



