METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Gemini 3.7 Flash Benchmark: Score 56, 1.7 Minutes per Task

Artificial Analysis measured 56 on its Intelligence Index and 1.7 minutes per task. It also led AA-AnalystAgent at 60%, ahead of Claude Opus 5.

Gemini 3.7 Flash Benchmark: Score 56, 1.7 Minutes per Task

Summary

  • Google DeepMind's Gemini 3.7 Flash scored 56 on the Artificial Analysis Intelligence Index at high reasoning effort, up 4 points from its predecessor, 3.6 Flash (52).
  • Average time per task fell from 2.0 minutes to 1.7 minutes, backed by an output speed of about 340 tokens per second, putting it on the "intelligence vs. time" Pareto frontier.
  • The list price stays unchanged at $1.50 input / $7.50 output per million tokens, but a discounted rate of $0.50 / $3.75 applies through the end of the year.

Google DeepMind's newly released Gemini 3.7 Flash completed an average task in 1.7 minutes in tests run by benchmarking firm Artificial Analysis. Its predecessor, Gemini 3.6 Flash, took 2.0 minutes on the same set of tasks. Getting smarter usually means getting slower — but this time the score went up while the time went down.

인공지능 모델별 작업당 시간과 지능 지수 비교 그래프와 산점도 차트
이미지: @ArtificialAnlys (X)

Up 4 points, down 18 seconds

The Artificial Analysis Intelligence Index combines multiple evaluations into a single score representing a model's overall capability. At high reasoning effort, Gemini 3.7 Flash scored 56, up 4 points from 3.6 Flash's 52. The speed gain comes from an output rate of roughly 340 tokens per second. As a result, the model now sits on the Pareto frontier of the "intelligence vs. time per task" chart — the line connecting models that are either the smartest for a given time budget or the fastest for a given intelligence level.

That said, it isn't free. At high reasoning effort, 3.7 Flash uses about 40% more tokens than its predecessor. Average output per task runs to 37k tokens, in the same range as Claude Fable 5 or Qwen3.8 Max. In other words, it thinks longer, but spits out that thinking much faster.

인공지능 모델별 작업당 비용과 지능 지수 비교 그래프와 비용 세부 내역 막대그래프
이미지: @ArtificialAnlys (X)

Three reasoning levels: which to choose

The new model offers three adjustable reasoning-effort settings: low, medium, and high. Lining up the numbers makes the trade-off clear.

SettingIntelligence IndexTime per task (min)Time comparison
Gemini 3.7 Flash (low)470.926
Gemini 3.7 Flash (medium)511.441
Gemini 3.7 Flash (high)561.750
Gemini 3.6 Flash522.059
Claude Opus 5 (max)638.0100

Even at the low setting, the model scores 47 — below predecessor 3.6 Flash's 52, but in less than half the time. At the other end, top-tier Claude Opus 5 (63) takes 8 minutes per task. Is a 7-point gain worth spending 4.7x more time? In practice, that's usually where the model-selection decision comes down.

인공지능 모델별 작업당 출력 토큰 수를 나타낸 막대그래프
이미지: @ArtificialAnlys (X)

List price unchanged, but a third of the cost through year-end

Pricing is also worth noting. The list price — $1.50 input / $7.50 output per million tokens — is identical to 3.6 Flash's. On top of that, Google is applying a discounted rate of $0.50 input / $3.75 output through the end of the year. Artificial Analysis calculated that, at the discounted rate, the cost per Intelligence Index task comes out to about $0.40.

Since the model uses roughly 40% more tokens, the effective cost at list price is actually higher than its predecessor's. That gap is currently masked by the limited-time discount running through year-end.

AA-Briefcase Elo와 GDPval-AA v2 지능 평가 점수 막대그래프
이미지: @ArtificialAnlys (X)

Tops the leaderboard in reading spreadsheets

The biggest score gap shows up in agentic evaluations. On AA-Briefcase, Artificial Analysis's own benchmark for knowledge-work tasks, 3.7 Flash (high) posted an Elo of 1132 — 169 points above its predecessor and edging out MiniMax-M3.

On AA-AnalystAgent, a newly introduced benchmark that has models answer complex questions about spreadsheets and documents, it topped the list with a pass^5 score of 60%. On the same benchmark, Claude Opus 5 (max) scored 54% and Claude Fable 5 scored 49%. A model that trails Opus 5 by 7 points on the overall Intelligence Index outperforms it on the specific task of reading a spreadsheet and finding the answer.

EvaluationGemini 3.7 Flash (high)Comparison models
Intelligence Index56Claude Opus 5: 63 · Grok 4.6: 61
AA-Briefcase (Elo)1132+169 vs. predecessor, just above MiniMax-M3
AA-AnalystAgent (pass^5)60%Claude Opus 5: 54% · Fable 5: 49%
AA-Analyst Agent와 AutomationBench-AA 점수 비교 막대그래프
이미지: @ArtificialAnlys (X)

Three Flash models in three months

Artificial Analysis noted this is the third new Flash model Google DeepMind has released in three months. In Google's Gemini lineup, the Flash line serves as the lighter, cheaper, faster "workhorse" tier below the flagship models — built for bulk processing and repeated agentic calls, where throughput per second matters more than a single flash of brilliance.

Competition in this segment has recently shifted from raw scores toward "time and cost per point of score." As this outlet covered on August 12, xAI's Grok 4.6 scored 61 on the Intelligence Index while keeping pricing unchanged from its predecessor, and Ant Group's Ling 3.0 Tiny scored 25 with just 1.3B active parameters, pushing the frontier of "small and cheap." 3.7 Flash sits in between — a mid-priced model targeting the time axis.

여러 지능 평가 항목별 인공지능 모델 점수 막대그래프 모음
이미지: @ArtificialAnlys (X)

What actually changes

For teams running agents on real workloads, wait time is cost. If you're summarizing twenty documents or digging through spreadsheets for answers hundreds of times a day, a 0.3-minute difference per task translates directly into throughput. 3.7 Flash makes a numbers-based case for replacing its predecessor on exactly that front.

The selection logic is also fairly simple now: use high reasoning effort for accuracy-critical analysis work, and drop to low for repetitive, lower-difficulty tasks like classification or extraction, to save time and tokens. That said, all these figures come from a third-party evaluator, Artificial Analysis, and real-world performance on your own data still needs to be verified directly. The fact that the discount only runs through year-end is also worth factoring into adoption timing.

Comments