All News
googlegeminigemini-3-8-flashreleasecodingcybersecurity

Gemini 3.8 Flash: Google's third Flash in six weeks, plus a defenders-only Cyber twin

Google released Gemini 3.8 Flash and a Cyber twin for vetted defenders on September 2. Third Flash in six weeks, same price as 3.7 — hungrier for tokens.

Vlad MakarovVlad Makarovreviewed and published
5 min read
Mentioned models
Gemini 3.8 Flash: Google's third Flash in six weeks, plus a defenders-only Cyber twin

Google released Gemini 3.8 Flash on September 2, three weeks after 3.7 Flash and as its third Flash-tier release in six weeks. Alongside it came Gemini 3.8 Flash Cyber, a cybersecurity variant sharing the same foundation but gated to vetted defenders through Google's new Fairwind Program. The standard model is live under the API id gemini-3.8-flash in AI Studio, the Gemini Enterprise Agent Platform, the Gemini app for AI Pro and Ultra subscribers, AI Mode in Search, and Google Sheets.

Three Flashes in six weeks

The cadence is the story Google chose to lead with: "our third Flash release in only six weeks," the launch post notes, positioning 3.8 Flash as "our best reasoning and coding model yet, at the same speed and low cost of 3.7." Both variants, per the blog, are "powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models." Tulsee Doshi and Raluca Ada Popa authored the post.

The release rhythm carries a quiet implication for the rest of the lineup. Ars Technica observes that Google has not shipped a frontier-level Gemini Pro since early 2026, and that each new Flash "makes it more likely that we'll never see the promised Gemini 3.5 Pro" — a model Google reportedly delayed when its coding performance could not match rivals.

Works harder, spends more tokens

The claimed gains come from a design choice the blog states plainly: "3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively." The same paragraph concedes the trade-off: "At times, the model might use more tokens to maximize performance, especially at higher effort levels."

Effort is now a dial. Developers chasing efficiency can run lower effort levels, or "continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads" — a vendor note that reads as an admission the improvement is bought with tokens rather than free. Google also credits part of the shared core's progress to "rigorous training in the highly demanding domain of cybersecurity," an unusual thing to say about a coding flagship.

Benchmarks

All figures below come from Google's launch post and developer docs — unaudited vendor numbers, pending independent evals.

BenchmarkGemini 3.8 FlashGemini 3.7 Flash
Terminal-Bench 2.190.8%81.6%
SWE-Bench Pro61.6%60.4%
SWE-Atlas51.9%48.0%
Humanity's Last Exam45.4%45.7%
HLE-Verified54.9%

The picture is uneven. Terminal-Bench 2.1 jumped more than nine points, while repository-level coding moved little (SWE-Bench Pro +1.2) and exam-style reasoning went nowhere (HLE −0.3). Google also claims the top of the DeepSWE v1.1 leaderboard for long-horizon engineering and wins on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark among the models it tested. Computer use remains the weak flank: Ars Technica notes 3.8 Flash improved on OSWorld-2.0 over 3.7 but stays far behind Claude Opus, "and GPT isn't great in this test, either."

The defenders-only Cyber twin

Gemini 3.8 Flash Cyber is the launch's second door, and everything about it is vendor-reported. Google says it reaches "frontier-level performance" on CyberGym, the standard vulnerability-discovery benchmark, surpassing 3.5 Flash Cyber and "significantly larger frontier models"; on an internal benchmark spanning 20 programming languages, it claims a vulnerability-discovery success rate above 70%. On CWE-Bench — an external patching benchmark run by Collinear — it sits on the Pareto frontier at 47.2% pass@1 versus a leading frontier model's 47.8%, at lower cost.

Google stresses the defensive bias: it prioritized vulnerability fixing "over offensive capabilities like exploitation," and claims its Chrome security team saw 2.6x more correct patches than from much larger commercial models, while its Cloud Vulnerability Research team found a critical vulnerability in under two hours — work that "usually takes months." Partner statements from Wiz and Palo Alto Networks attest to the capability; those are endorsements from Google's own ecosystem, not independent measurement. DataCamp's review lands on the same caution: "almost all the data so far comes from Google itself."

The gating is real either way. Cyber ships with a "more permissive set of mitigations" than the base model and reaches only trusted government authorities, critical-infrastructure operators, and software maintainers through Fairwind — per Ars Technica, "trusted testers and governments." It replaces the 3.5 Flash Cyber tier. What the model can and cannot do outside Google's own evals is exactly what the program's secrecy makes hardest to check.

Same price, cheaper than the flagships

Pricing matches 3.7 Flash: an introductory $0.75 per million input tokens and $3.75 per million output through December 31, 2026, then $1.50/$7.50. That undercuts the reasoning flagships by an order of magnitude — DataCamp's comparison puts GPT-5.6 Sol at roughly $5/$30 and Claude Fable 5.1 at $10/$50, the pricing race Anthropic entered with its Fable 5.1 / Mythos 5.1 launch. Ars Technica reads the teaser rate as a necessity: rival labs keep cutting token prices to hold business attention, and "it's likely there will be new models available long before the price changes."

The bottom line

Reception on r/singularity has been engaged but not ecstatic — the benchmark-chart thread drew roughly 780 points and more than 200 comments, mixing impressed replies with the familiar "third Flash in six weeks, where is Pro?" grumbling. We track Gemini 3.8 Flash with benchmarks on the model page.

The honest read: 3.8 Flash looks like a genuine coding upgrade over a model released three weeks earlier, at an unchanged price, and the Cyber twin is an unusually disciplined product decision — patching over exploitation, defenders over everyone. But the flagship-grade claims rest on Google's own charts, and the Cyber numbers rest on Google's own benchmarks plus partner endorsements. What would settle it: independent runs of Terminal-Bench and CWE-Bench, and a third-party look at whether "works harder" survives contact with a production token budget.

Related Articles

Scroll down

to load the next article