
Summary
- Google DeepMind released the Gemini 3.8 Flash model card on September 2, 2026, built on top of Gemini 3.7 Flash
- It accepts text, image, audio, and video input, supports a context window of up to 1 million tokens, and can output up to 64,000 tokens
- It's rolling out across six channels — including the Gemini app, Google AI Studio, and the Gemini API — and is tuned for software engineering and agentic work
Google DeepMind published the model card for Gemini 3.8 Flash on September 2, 2026. It's the latest iteration in the Gemini 3 line, and the headline change is improved performance on software engineering and agentic knowledge work.
Google DeepMind official website
To put it plainly: Gemini Flash is Google's experimental lineup, refined with point releases every few weeks.
From 3.6 to 3.8 in three weeks
Last month, on the 12th, TestingCatalog spotted a "Gemini 3.6 Flash" test banner on the Gemini Enterprise screen. Less than three weeks later, Google has formally posted the 3.8 Flash model card. The card points readers to the 3.7 Flash model card for reference, noting that the two models share the same architecture, training data, hardware, and software.
What actually changed
The model card boils the changes down to two things. First, better performance on software engineering and agentic knowledge work. Second, continued support for effort levels, which let users balance quality, cost, and latency. The specific numbers are laid out in the official benchmark section below.
Input covers text, image, audio, and video files, with a context window of up to 1 million tokens. Output tops out at 64,000 tokens of text. These figures match 3.7 Flash exactly, which suggests this update is less about a new architecture and more about performance tuning within the same framework. The card points readers to the evaluation methodology page for details on how the tests were run.
| Item | Detail |
|---|---|
| Base model | Gemini 3.7 Flash |
| Input | Text, image, audio, video |
| Context window | Up to 1 million tokens |
| Output | Text, up to 64,000 tokens |
| Primary use cases | Software engineering, agentic tasks, complex knowledge work |
The official benchmark table — 5 to 7 times cheaper, but wins and big losses split down the middle
Page 5 of the model card carries a 14-benchmark comparison table. The table is embedded as an image inside the PDF, so it won't show up if you only scrape the body text. Below is that table reproduced in full.
Start with pricing. 3.8 Flash runs $0.75 per million input tokens and $3.75 per million output tokens — about a seventh the cost of Claude Opus 5 ($5.00 / $25.00). That said, this is an introductory rate that only holds through December 31, 2026. Starting January 1, 2027, it rises to $1.50 input / $7.50 output.
| Per million tokens | 3.8 Flash | 3.7 Flash | Opus 5 | Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Input | $0.75 (list price $1.50) | $0.75 | $5.00 | $2.00 | $4.00 | $2.00 |
| Output | $3.75 (list price $7.50) | $3.75 | $25.00 | $10.00 | $20.00 | $12.00 |
Where 3.8 Flash ranks first
| Benchmark | 3.8 Flash | 3.7 Flash | Opus 5 | Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Vals Finance Agent v2 (financial analysis) | 61.4% | 59.0% | 58.6% | 53.9% | 53.8% | 54.4% |
| Harvey legal agent | 10.0% | 8.8% | 6.7% | 5.0% | 2.5% | 0.8% |
| Terminal-bench 2.1 (terminal coding) | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
| CharXiv Reasoning (complex chart interpretation) | 86.2% | 84.5% | 83.7% | 70.1% | 85.8% | 85.9% |
| LVBench (long-video understanding) | 87.8% | 85.4% | 75.4% | 68.5% | 82.1% | 78.9% |
| HLE-Verified (expert-knowledge reasoning) | 54.9% | 53.6% | 54.4% | 31.0% | 54.5% | 51.1% |
| BioMysteryBench (hard difficulty) | 56.5% | 43.5% | 49.4% | 34.1% | 44.7% | 49.4% |
| LABBench2 (biology research) | 86.2% | 82.1% | 84.2% | 80.1% | 82.1% | 81.2% |
Where 3.8 Flash loses — and this might matter more
| Benchmark | 3.8 Flash | Top model |
|---|---|---|
| Terminal-bench 4.0 (general-purpose agent) | 19.1% | Opus 5: 51.8% |
| OSWorld-2.0 (computer operation) | 59.0% | Opus 5: 75.4% |
| GDPVal-AA v2 (knowledge work, Elo) | 1545 | Opus 5: 1824 |
| DeepSWE v1.1 (long-horizon software engineering) | 73.7% | Opus 5: 74.0% |
| GDP.PDF (professional PDF comprehension) | 35.0% | GPT-5.6 Sol: 40.0% |
The table Google put front and center in its promotional materials is the one above. Set it alongside the one below, and the picture changes. On general-purpose agent work (Terminal-bench 4.0), it's 19.1% versus 51.8%, and on computer operation (OSWorld-2.0), it's 59.0% versus 75.4% — both large losses to Opus 5. In other words, 3.8 Flash is cheap and strong on narrow tasks with defined tools, but tasks requiring the model to figure things out in unfamiliar environments still belong to frontier models.
Customer quotes DeepMind posted on its official product page point in the same direction. Tai Tran, head of AI product at document-search company Glean, said the model "completed more than three times the work of 3.7 Flash on long, document-heavy workflows," while Haoyuan Guo, CTO of game-development agent maker Loopit, said his team uses it for coding, asset reasoning, and visual verification. Both are quotes Google chose to feature, so read them with that in mind.

What the model card quietly notes
Beyond the numbers, there are three things worth flagging in the card.
First, multilingual safety slipped 5.4 points. In Google's automated safety evaluations, text safety improved 0.4 points and image safety was unchanged, but multilingual safety alone got worse. The card states plainly that "non-English safety performance regressed slightly compared to 3.7 Flash." That's not a detail Korean-language users should brush past.
Second, frontier safety evaluation was not run on 3.8 itself. The card notes that 3.7 Flash was evaluated and did not reach risk thresholds (T/CCL), and that 3.8 is assumed to reach the same conclusion "since it does not introduce meaningful new capabilities or a substantial performance increase."
Third, the knowledge cutoff is March 2026, though the card flags that some domains may still reflect a January 2025 level of knowledge. Known limitations listed include hallucination, jailbreak resistance, occasional latency or timeouts, and excessive token usage at high effort levels.
The official product page shipped with four demos: a 3D puzzle game built with a single line of repeated instructions in Antigravity; the DOS-style Google Maps shown above; a terrain cross-section built from real USGS and OpenWeather data; and a 3D viewer that renders a hardware teardown using Three.js.
Where you can use it
Gemini 3.8 Flash is available through six channels: the Gemini app, the Gemini Enterprise agent platform, Google AI Studio, the Gemini API, Google AI Mode, and Google Antigravity.
Developers can select the model in Google AI Studio or the Gemini API by specifying gemini-3.8-flash and choosing a quality/cost/latency combination via the effort level parameter. Everyday users can access it through the model-selection menu in the Gemini app.
| Channel | Description |
|---|---|
| Gemini app | Consumer-facing chatbot |
| Gemini Enterprise Agent Platform | Enterprise agent platform |
| Google AI Studio | Developer prototyping tool |
| Gemini API | API for service integration |
| Google AI Mode | AI answer mode within search |
| Google Antigravity | Agentic development platform |
Limitations and safeguards
The model card acknowledges that hallucination — stating plausible falsehoods as fact — remains an unresolved issue. Google says it continues to improve jailbreak resistance and has recently tightened its overall frontier safety mitigations.
In Korea
In Korea, Google formally launched its personal AI agent "Gemini Spark" on July 30, and this month it reportedly relaunched two promotions — a free paid Gemini membership for university students and a discount on the AI Pro plan — according to local reporting. Given that 3.8 Flash is tuned for software engineering and agentic tasks, it's likely to underpin these domestic agent services.
A second model, released the same day — 3.8 Flash Cyber
Google DeepMind announced "two new Gemini models" the same day on its official X account. One is the 3.8 Flash covered above. The other is 3.8 Flash Cyber, a dedicated model built to find security vulnerabilities and generate patches for them autonomously.
The problem Google frames it around: security teams currently have to choose between "large, expensive, slow frontier models" and "cheap models that can't handle complex code repairs." Google says 3.8 Flash Cyber generates fixes deployable within minutes, inside an organization's own cloud environment.

CyberGym — claiming the top spot in C/C++ vulnerability discovery

| CyberGym Pass@1 (C/C++ vulnerability discovery) | Score |
|---|---|
| Gemini 3.8 Flash Cyber | 86.2% |
| GPT-5.5-Cyber | 85.6% |
| Mythos 5 | 83.8% |
| GPT-5.6 Sol | 83.6% |
| Gemini 3.5 Flash Cyber | 77.5% |
It comes in first, but the margin over second-place GPT-5.5-Cyber (85.6%) is just 0.6 points. What stands out more is the jump from its own predecessor, Gemini 3.5 Flash Cyber (77.5%) — up 8.7 points.
CWE-Bench — similar accuracy, a third of the cost

The horizontal axis shows average cost per rollout (cheaper toward the right); the vertical axis shows accuracy.
| CWE-Bench Pass@1 | Accuracy | Approx. cost per rollout |
|---|---|---|
| Claude Fable 5 | 47.9% | ~$10.3 |
| Gemini 3.8 Flash Cyber | 47.3% | ~$3.6 |
| GPT-5.6 Sol | 44.3% | ~$2.4 |
| Gemini 3.7 Flash | 44.2% | ~$1.5 |
| Opus 4.8 | 42.2% | ~$2.5 |
| Grok 4.6 | 38.4% | ~$2.0 |
| Hy4 Preview | 33.8% | ~$0.8 |
| DeepSeek V4 Flash | 30.5% | ~$0.2 |
Top accuracy goes to Claude Fable 5 (47.9%). 3.8 Flash Cyber trails by just 0.6 points at 47.3%, but does it at roughly a third of the cost — which is exactly why Google plotted this as a Pareto curve. The point of the chart isn't "we're the best." It's "we get nearly the same result for far less money."
Google also shared one figure drawn from real code rather than benchmarks: testing against Chrome codebases, it says the model produced 2.6 times more valid fixes. The post doesn't specify what model it was compared against.
Not for everyone — the Fairwind Program
3.8 Flash Cyber isn't broadly available. Google says it's granting early trusted access through a separate channel called the Fairwind Program, aimed at national cybersecurity authorities and operators of essential services like telecommunications and energy. The stated rationale is protecting critical public infrastructure.
The ability to both find vulnerabilities and automatically generate patches cuts both ways — it's just as useful to attackers as to defenders. Restricting access to government bodies and critical industries reads as an acknowledgment from Google that it's aware of that double edge. Meanwhile, the standard 3.8 Flash is now rolling out through Google Antigravity and the AI Studio API.
Editor's take
It's no coincidence that "software engineering" and "agentic knowledge work" keep coming up. In the same window, Gemini cut token consumption by up to 88% through agentic video understanding, and was reportedly testing a toggle between chat and task modes in its enterprise interface. All three moves point in one direction: shifting weight from a chatbot that answers well to an agent that finishes the work for you.
Compared with previous generations, what stands out isn't the numbers — it's the deployment logic. The fact that architecture, training data, and even hardware carry over unchanged from 3.7 Flash suggests Google is choosing to iterate within the same framework every few weeks, rather than rebuilding the model from scratch each time. Put these point releases into practice, and the takeaway is usually the same: expect small gains on specific tasks — this time coding and agentic work — rather than a big leap in overall capability.
For companies in Korea, the practical move isn't deciding whether to adopt this model outright, but tweaking the effort level parameter in an existing Gemini API pipeline to rebalance cost and latency. Since detailed benchmark comparisons aren't fully public yet, it's too early to gauge the exact accuracy gains — the only real way to know is to run it against actual workloads.
In the coming weeks, expect agentic processing to expand further into the Gemini app and YouTube's "Ask YouTube" feature, with 3.8 Flash likely serving as the underlying model in some of those cases. With point releases moving this fast, it wouldn't be surprising if the next one arrives in days rather than weeks.





Comments