One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Gemini 3.7 Flash gains 4 intelligence points, cuts response time

Scores 56 on Artificial Analysis Intelligence Index at 1.7 minutes per task — third Flash model in three months

AI 모델별 지능 지수와 처리 시간을 비교한 벤치마크 차트

이미지: X — 벤치마크·평가 화면 갈무리

Summary

  • Google DeepMind's Gemini 3.7 Flash scored 56 on the Artificial Analysis Intelligence Index at high reasoning effort, up 4 points from its predecessor 3.6 Flash's 52.
  • Average time per task dropped from 2.0 minutes to 1.7 minutes, backed by an output speed of about 340 tokens per second, placing it on the "intelligence vs. time" Pareto frontier.
  • List price stays frozen at $1.50 per million input tokens and $7.50 per million output tokens, but a discounted rate of $0.50/$3.75 applies through year-end.
모델
Gemini 3.7 Flash (개발: Google DeepMind)
Intelligence Index
high 56점 · medium 51점 · low 47점 (전작 3.6 Flash 52점)
과제당 시간
high 1.7분 · medium 1.4분 · low 0.9분 (3.6 Flash는 2.0분)
출력 속도
초당 약 340토큰
가격
정가 100만 토큰당 입력 1.50달러 / 출력 7.50달러, 연말까지 할인가 0.50달러 / 3.75달러
토큰 사용량
high 기준 과제당 평균 37k 출력 토큰, 전작 대비 약 40% 증가
AA-Briefcase
1132 Elo, 전작 대비 +169로 MiniMax-M3 바로 위
AA-AnalystAgent
pass^5 60%로 1위 (Claude Opus 5 max 54%, Claude Fable 5 49%)

Google DeepMind's newly released Gemini 3.7 Flash completed an average task in 1.7 minutes, according to measurements by benchmarking firm Artificial Analysis. Its predecessor, Gemini 3.6 Flash, took 2.0 minutes on the same set of tasks. Typically, getting smarter means getting slower — but this time, the score went up while the time went down.

이미지: X — 벤치마크·평가

4 points up, 18 seconds faster

The Artificial Analysis Intelligence Index combines multiple evaluations into a single composite score of a model's overall capability. Gemini 3.7 Flash scored 56 at high reasoning effort, up 4 points from 3.6 Flash's 52. The speed gain comes from an output rate of about 340 tokens per second. As a result, the model now sits on the Pareto frontier of the "intelligence vs. time per task" chart — the line connecting models that are either the smartest at a given time budget or the fastest at a given intelligence level.

That said, it isn't free. At high reasoning effort, 3.7 Flash uses roughly 40% more tokens than its predecessor. Average output per task runs to 37k tokens, on par with Claude Fable 5 or Qwen3.8 Max. In other words, it thinks longer, but spits out that thinking much faster.

Gemini 3.7 Flash benchmark chart
Intelligence Index vs. Time per Task · Artificial Analysis
이미지: X — 벤치마크·평가

Three reasoning tiers: which to choose

The new model offers three reasoning-effort levels — low, medium, and high. Laid side by side, the tradeoffs become clear.

SettingIntelligence IndexTime per task (min)Time comparison
Gemini 3.7 Flash (low)470.926
Gemini 3.7 Flash (medium)511.441
Gemini 3.7 Flash (high)561.750
Gemini 3.6 Flash522.059
Claude Opus 5 (max)638.0100

Even at the low setting, the model scores 47 — short of predecessor 3.6 Flash's 52, but in less than half the time. At the other end, top-tier Claude Opus 5 (63) takes 8 minutes per task. Is a 7-point gap worth spending 4.7 times as long? That's usually where the real-world model-selection decision lands.

이미지: X — 벤치마크·평가

List price unchanged, but a third of that through year-end

Pricing is also worth noting. The list price — $1.50 per million input tokens and $7.50 per million output tokens — matches 3.6 Flash exactly. On top of that, Google is applying a discounted rate of $0.50 input / $3.75 output through the end of the year. Artificial Analysis notes that at the discounted rate, cost per Intelligence Index task comes out to roughly $0.40.

Since the model uses about 40% more tokens, its real-world cost at list price is actually higher than its predecessor's. That gap is currently masked by the temporary year-end discount.

이미지: X — 벤치마크·평가

Top of the list on reading spreadsheets

The biggest gains showed up in agentic evaluations. On AA-Briefcase, Artificial Analysis's own benchmark for knowledge-work tasks, 3.7 Flash (high) posted an Elo of 1132 — 169 points above its predecessor and narrowly ahead of MiniMax-M3.

On AA-AnalystAgent, a newly introduced benchmark that has models answer complex questions about spreadsheets and documents, 3.7 Flash topped the list with a pass^5 score of 60%. On the same benchmark, Claude Opus 5 (max) scored 54% and Claude Fable 5 scored 49%. So a model that trails Opus 5 by 7 points on the overall Intelligence Index actually beats it on the specific task of reading a spreadsheet and finding the answer.

EvaluationGemini 3.7 Flash (high)Comparison models
Intelligence Index56Claude Opus 5: 63 · Grok 4.6: 61
AA-Briefcase (Elo)1132+169 vs. predecessor, just above MiniMax-M3
AA-AnalystAgent (pass^5)60%Claude Opus 5: 54% · Fable 5: 49%
이미지: X — 벤치마크·평가

Third Flash release in three months

Artificial Analysis noted that this marks the third new Flash model Google DeepMind has released in three months. Within Google's Gemini lineup, the Flash line serves as the lighter, cheaper, faster "workhorse" tier beneath the flagship models — built for high-volume processing and repeated agentic calls, where throughput per second matters more than any single flash of brilliance.

Competition in this segment has recently shifted from raw scores toward "time and cost per score point." As we reported on August 12, xAI's Grok 4.6 scored 61 on the Intelligence Index while keeping pricing unchanged from its predecessor, and Ant Group's Ling 3.0 Tiny scored 25 with just 1.3B active parameters, pushing the frontier on the "small and cheap" end. 3.7 Flash sits in between — a mid-priced model targeting the time axis.

이미지: X — 벤치마크·평가

What actually changes

For teams running agents on real workloads, latency is cost. If you're summarizing twenty documents or digging through spreadsheets for answers hundreds of times a day, a 0.3-minute difference per task directly translates into a difference in throughput. On that front, 3.7 Flash makes a numerical case for replacing its predecessor.

The choice itself is also fairly simple now: use high effort for accuracy-critical analysis work, and dial down to low for repetitive, low-difficulty tasks like classification or extraction, to save time and tokens. That said, all of these figures come from a single third-party evaluator, Artificial Analysis, and real performance on your own workloads still needs to be tested directly. The fact that the discounted pricing only runs through year-end is also a factor worth weighing when deciding when to adopt it.