METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Grok 4.6 scores 61 on intelligence index, matches GPT-5.6 at the same price

Score jumps 5 points a month after launch, takes the runner-up spot to Claude Opus 5 in agentic tasks

Grok 4.6 scores 61 on intelligence index, matches GPT-5.6 at the same price

Summary

  • xAI's Grok 4.6 (listed as "SpaceXAI's Grok 4.6") scored 61 on the Artificial Analysis Intelligence Index, putting it in the same top tier as GPT-5.6 Sol
  • Across three practical benchmarks — GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1 — it matched or beat every model except Claude Opus 5
  • Pricing stayed the same as the previous Grok 4.5, and per-task cost is on par with Kimi K3, far below Claude Opus 5 and GPT-5.6 Sol

Five points in a month — Grok's rapid comeback

xAI's large language model Grok 4.6 scored 61 on the Intelligence Index from evaluator Artificial Analysis. The post in question labeled the model "SpaceXAI's Grok 4.6." A score of 61 puts it on par with GPT-5.6 Sol, placing it among the top-tier models. What stands out is the speed: the score jumped 5 points in just over a month since the previous Grok 4.5 was released.

여러 AI 모델의 성능 점수를 비교한 세 개의 가로 막대그래프 차트
이미지: @ArtificialAnlys (X)

Booking the runner-up spot behind Claude Opus 5

Grok 4.6 showed strength in practical, work-oriented tasks such as knowledge work, terminal operation, and customer service. Across three benchmarks measured independently by Artificial Analysis — GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1 — Grok 4.6 matched or outperformed every model except Claude Opus 5.

ModelGDPval-AA v2 (Elo)τ³-BankingTerminal-Bench v2.1
Claude Opus 5 (max)184942.1%89.1%
Grok 4.6 (high)175350.7%88.4%
Claude Fable 5174138.1%84.6%
Qwen3.8 Max173751.3%81.3%
GPT-5.6 Sol (max)172844.3%88.0%
Kimi K3 (max)168246.0%85.0%

GDPval-AA v2 grades real-world work tasks against a human baseline (1000 points). τ³-Banking measures banking-scenario tasks, and Terminal-Bench v2.1 measures the ability to operate a computer terminal. Grok 4.6 held the No. 2 spot across all three metrics.

인공지능 지능지수와 작업당 비용을 비교한 산점도 그래프
이미지: @ArtificialAnlys (X)

Cheap and capable — the cost-performance angle

Grok 4.6's standard pricing is unchanged from Grok 4.5. Yet its intelligence score rose 5 points. According to Artificial Analysis, its per-task cost is comparable to Kimi K3 and far lower than Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5. Artificial Analysis assessed that when performance and cost are plotted together, the model lands on the Pareto frontier.

The model also made its first appearance on AA-Briefcase, a private benchmark covering long-duration knowledge work tasks. It scored an Elo of 1577 there, again ranking second behind Claude Opus 5. Artificial Analysis noted that it "consistently performed strongly across grading criteria, document structure, and analysis quality."

AA-Briefcase Elo 점수를 비교한 가로 막대그래프 차트
이미지: @ArtificialAnlys (X)

What the Intelligence Index means

The Intelligence Index combines scores from multiple benchmarks into a single ranking of models. Meta's Muse Spark 1.2, released on August 8, previously scored 54 on this index, tying for third place with SpaceXAI. By contrast, Ant Group's open-weight small model Ling 3.0 Tiny, released August 12, scored only 25. Seen in this context, Grok 4.6's score of 61 places it in the crowded top tier dominated by large closed models, while the gap with open-weight small models remains wide.

다양한 AI 평가 항목별 점수를 비교한 여러 가로 막대그래프 차트
이미지: @ArtificialAnlys (X)

So what changes

Until now, top-tier intelligence and low cost were seen as difficult to achieve together. Grok 4.6 kept its price unchanged while raising performance, delivering both at once. For enterprises that need to run long agentic workloads repeatedly, this adds another option that offers top-tier performance without the cost burden — even if it falls short of Claude Opus 5's absolute best quality. It can be read as a sign that the axis of frontier model competition is shifting from "who is smartest" to "who can do more at the same cost."

Comments