METAL LAB

Google unveils Gemini 3.8 Flash, tuned for coding and agentic tasks

The new model builds on Gemini 3.7 Flash, ships across six deployment channels, and supports up to 1 million tokens of context

Summary

  • Google DeepMind released the Gemini 3.8 Flash model card on September 2, 2026, built on top of Gemini 3.7 Flash
  • It accepts text, image, audio, and video input, supports a context window of up to 1 million tokens, and can output up to 64,000 tokens
  • It's rolling out across six channels — including the Gemini app, Google AI Studio, and the Gemini API — and is tuned for software engineering and agentic work

Google DeepMind published the model card for Gemini 3.8 Flash on September 2, 2026. It's the latest iteration in the Gemini 3 line, and the headline change is improved performance on software engineering and agentic knowledge work.

A drawing shows 3.8 Flash inheriting the same dotted-seed skeleton as 3.7 Flash, tuned up for coding and agentic tasks and rendered as growing dots, with light rays symbolizing the model spreading simultaneously across six deployment channels.A drawing shows 3.8 Flash inheriting the same dotted-seed skeleton as 3.7 Flash, tuned up for coding and agentic tasks and rendered as growing dots, with light rays symbolizing the model spreading simultaneously across six deployment channels.

Google DeepMind official website

To put it plainly: Gemini Flash is Google's experimental lineup, refined with point releases every few weeks.

From 3.6 to 3.8 in three weeks

Last month, on the 12th, TestingCatalog spotted a "Gemini 3.6 Flash" test banner on the Gemini Enterprise screen. Less than three weeks later, Google has formally posted the 3.8 Flash model card. The card points readers to the 3.7 Flash model card for reference, noting that the two models share the same architecture, training data, hardware, and software.

What actually changed

The model card boils the changes down to two things. First, better performance on software engineering and agentic knowledge work. Second, continued support for effort levels, which let users balance quality, cost, and latency. The specific numbers are laid out in the official benchmark section below.

Input covers text, image, audio, and video files, with a context window of up to 1 million tokens. Output tops out at 64,000 tokens of text. These figures match 3.7 Flash exactly, which suggests this update is less about a new architecture and more about performance tuning within the same framework. The card points readers to the evaluation methodology page for details on how the tests were run.

ItemDetail
Base modelGemini 3.7 Flash
InputText, image, audio, video
Context windowUp to 1 million tokens
OutputText, up to 64,000 tokens
Primary use casesSoftware engineering, agentic tasks, complex knowledge work

The official benchmark table — 5 to 7 times cheaper, but wins and big losses split down the middle

Page 5 of the model card carries a 14-benchmark comparison table. The table is embedded as an image inside the PDF, so it won't show up if you only scrape the body text. Below is that table reproduced in full.

Start with pricing. 3.8 Flash runs $0.75 per million input tokens and $3.75 per million output tokens — about a seventh the cost of Claude Opus 5 ($5.00 / $25.00). That said, this is an introductory rate that only holds through December 31, 2026. Starting January 1, 2027, it rises to $1.50 input / $7.50 output.

Per million tokens3.8 Flash3.7 FlashOpus 5Sonnet 5GPT-5.6 SolGPT-5.6 Terra
Input$0.75 (list price $1.50)$0.75$5.00$2.00$4.00$2.00
Output$3.75 (list price $7.50)$3.75$25.00$10.00$20.00$12.00

Where 3.8 Flash ranks first

Benchmark3.8 Flash3.7 FlashOpus 5Sonnet 5GPT-5.6 SolGPT-5.6 Terra
Vals Finance Agent v2 (financial analysis)61.4%59.0%58.6%53.9%53.8%54.4%
Harvey legal agent10.0%8.8%6.7%5.0%2.5%0.8%
Terminal-bench 2.1 (terminal coding)89.4%85.8%89.1%80.4%88.8%87.4%
CharXiv Reasoning (complex chart interpretation)86.2%84.5%83.7%70.1%85.8%85.9%
LVBench (long-video understanding)87.8%85.4%75.4%68.5%82.1%78.9%
HLE-Verified (expert-knowledge reasoning)54.9%53.6%54.4%31.0%54.5%51.1%
BioMysteryBench (hard difficulty)56.5%43.5%49.4%34.1%44.7%49.4%
LABBench2 (biology research)86.2%82.1%84.2%80.1%82.1%81.2%

Where 3.8 Flash loses — and this might matter more

Benchmark3.8 FlashTop model
Terminal-bench 4.0 (general-purpose agent)19.1%Opus 5: 51.8%
OSWorld-2.0 (computer operation)59.0%Opus 5: 75.4%
GDPVal-AA v2 (knowledge work, Elo)1545Opus 5: 1824
DeepSWE v1.1 (long-horizon software engineering)73.7%Opus 5: 74.0%
GDP.PDF (professional PDF comprehension)35.0%GPT-5.6 Sol: 40.0%

The table Google put front and center in its promotional materials is the one above. Set it alongside the one below, and the picture changes. On general-purpose agent work (Terminal-bench 4.0), it's 19.1% versus 51.8%, and on computer operation (OSWorld-2.0), it's 59.0% versus 75.4% — both large losses to Opus 5. In other words, 3.8 Flash is cheap and strong on narrow tasks with defined tools, but tasks requiring the model to figure things out in unfamiliar environments still belong to frontier models.

Customer quotes DeepMind posted on its official product page point in the same direction. Tai Tran, head of AI product at document-search company Glean, said the model "completed more than three times the work of 3.7 Flash on long, document-heavy workflows," while Haoyuan Guo, CTO of game-development agent maker Loopit, said his team uses it for coding, asset reasoning, and visual verification. Both are quotes Google chose to feature, so read them with that in mind.

wWEwX2i0HtKcCbqllmsPq68sJY8vl0qTP1O8gXfEl81ulUVafxvPrgtY p60TZGOMTgiz 5u VQvGFYQckQVol37BnvcPoyXENT657y PyMsRE5oEw=w1440 rw lo

What the model card quietly notes

Beyond the numbers, there are three things worth flagging in the card.

First, multilingual safety slipped 5.4 points. In Google's automated safety evaluations, text safety improved 0.4 points and image safety was unchanged, but multilingual safety alone got worse. The card states plainly that "non-English safety performance regressed slightly compared to 3.7 Flash." That's not a detail Korean-language users should brush past.

Second, frontier safety evaluation was not run on 3.8 itself. The card notes that 3.7 Flash was evaluated and did not reach risk thresholds (T/CCL), and that 3.8 is assumed to reach the same conclusion "since it does not introduce meaningful new capabilities or a substantial performance increase."

Third, the knowledge cutoff is March 2026, though the card flags that some domains may still reflect a January 2025 level of knowledge. Known limitations listed include hallucination, jailbreak resistance, occasional latency or timeouts, and excessive token usage at high effort levels.

The official product page shipped with four demos: a 3D puzzle game built with a single line of repeated instructions in Antigravity; the DOS-style Google Maps shown above; a terrain cross-section built from real USGS and OpenWeather data; and a 3D viewer that renders a hardware teardown using Three.js.

Where you can use it

Gemini 3.8 Flash is available through six channels: the Gemini app, the Gemini Enterprise agent platform, Google AI Studio, the Gemini API, Google AI Mode, and Google Antigravity.

Developers can select the model in Google AI Studio or the Gemini API by specifying gemini-3.8-flash and choosing a quality/cost/latency combination via the effort level parameter. Everyday users can access it through the model-selection menu in the Gemini app.

ChannelDescription
Gemini appConsumer-facing chatbot
Gemini Enterprise Agent PlatformEnterprise agent platform
Google AI StudioDeveloper prototyping tool
Gemini APIAPI for service integration
Google AI ModeAI answer mode within search
Google AntigravityAgentic development platform

Limitations and safeguards

The model card acknowledges that hallucination — stating plausible falsehoods as fact — remains an unresolved issue. Google says it continues to improve jailbreak resistance and has recently tightened its overall frontier safety mitigations.

In Korea

In Korea, Google formally launched its personal AI agent "Gemini Spark" on July 30, and this month it reportedly relaunched two promotions — a free paid Gemini membership for university students and a discount on the AI Pro plan — according to local reporting. Given that 3.8 Flash is tuned for software engineering and agentic tasks, it's likely to underpin these domestic agent services.

A second model, released the same day — 3.8 Flash Cyber

Google DeepMind announced "two new Gemini models" the same day on its official X account. One is the 3.8 Flash covered above. The other is 3.8 Flash Cyber, a dedicated model built to find security vulnerabilities and generate patches for them autonomously.

The problem Google frames it around: security teams currently have to choose between "large, expensive, slow frontier models" and "cheap models that can't handle complex code repairs." Google says 3.8 Flash Cyber generates fixes deployable within minutes, inside an organization's own cloud environment.

31jyeMWAamsR699oN thAvrro WbOIIQ5R r2jaYhC5Fb7pcP7eQNl3cRVriT4rNg2rdEEeDtOGBGuRZy0O9TaaCsuhSxpxkJ0vunrgnseAMN8mm=w1440 h810 n nu rw lo

CyberGym — claiming the top spot in C/C++ vulnerability discovery

CyberGym Pass@1 benchmark — Gemini 3.8 Flash Cyber leads at 86.2%
CyberGym Pass@1 (C/C++ vulnerability discovery)Score
Gemini 3.8 Flash Cyber86.2%
GPT-5.5-Cyber85.6%
Mythos 583.8%
GPT-5.6 Sol83.6%
Gemini 3.5 Flash Cyber77.5%

It comes in first, but the margin over second-place GPT-5.5-Cyber (85.6%) is just 0.6 points. What stands out more is the jump from its own predecessor, Gemini 3.5 Flash Cyber (77.5%) — up 8.7 points.

CWE-Bench — similar accuracy, a third of the cost

CWE-Bench Pass@1 versus per-rollout cost — 3.8 Flash Cyber sits above the Pareto curve

The horizontal axis shows average cost per rollout (cheaper toward the right); the vertical axis shows accuracy.

CWE-Bench Pass@1AccuracyApprox. cost per rollout
Claude Fable 547.9%~$10.3
Gemini 3.8 Flash Cyber47.3%~$3.6
GPT-5.6 Sol44.3%~$2.4
Gemini 3.7 Flash44.2%~$1.5
Opus 4.842.2%~$2.5
Grok 4.638.4%~$2.0
Hy4 Preview33.8%~$0.8
DeepSeek V4 Flash30.5%~$0.2

Top accuracy goes to Claude Fable 5 (47.9%). 3.8 Flash Cyber trails by just 0.6 points at 47.3%, but does it at roughly a third of the cost — which is exactly why Google plotted this as a Pareto curve. The point of the chart isn't "we're the best." It's "we get nearly the same result for far less money."

Google also shared one figure drawn from real code rather than benchmarks: testing against Chrome codebases, it says the model produced 2.6 times more valid fixes. The post doesn't specify what model it was compared against.

Not for everyone — the Fairwind Program

3.8 Flash Cyber isn't broadly available. Google says it's granting early trusted access through a separate channel called the Fairwind Program, aimed at national cybersecurity authorities and operators of essential services like telecommunications and energy. The stated rationale is protecting critical public infrastructure.

The ability to both find vulnerabilities and automatically generate patches cuts both ways — it's just as useful to attackers as to defenders. Restricting access to government bodies and critical industries reads as an acknowledgment from Google that it's aware of that double edge. Meanwhile, the standard 3.8 Flash is now rolling out through Google Antigravity and the AI Studio API.

Editor's take

It's no coincidence that "software engineering" and "agentic knowledge work" keep coming up. In the same window, Gemini cut token consumption by up to 88% through agentic video understanding, and was reportedly testing a toggle between chat and task modes in its enterprise interface. All three moves point in one direction: shifting weight from a chatbot that answers well to an agent that finishes the work for you.

Compared with previous generations, what stands out isn't the numbers — it's the deployment logic. The fact that architecture, training data, and even hardware carry over unchanged from 3.7 Flash suggests Google is choosing to iterate within the same framework every few weeks, rather than rebuilding the model from scratch each time. Put these point releases into practice, and the takeaway is usually the same: expect small gains on specific tasks — this time coding and agentic work — rather than a big leap in overall capability.

For companies in Korea, the practical move isn't deciding whether to adopt this model outright, but tweaking the effort level parameter in an existing Gemini API pipeline to rebalance cost and latency. Since detailed benchmark comparisons aren't fully public yet, it's too early to gauge the exact accuracy gains — the only real way to know is to run it against actual workloads.

In the coming weeks, expect agentic processing to expand further into the Gemini app and YouTube's "Ask YouTube" feature, with 3.8 Flash likely serving as the underlying model in some of those cases. With point releases moving this fast, it wouldn't be surprising if the next one arrives in days rather than weeks.

Comments