Gemini 3.8 Flash: Pricing, Migration and Flash Cyber
Gemini 3.8 Flash costs $0.75/$3.75 per million tokens until Dec 31, 2026, then doubles. Here's what changes vs 3.7 Flash and what Flash Cyber adds.
Google launched Gemini 3.8 Flash on September 2, 2026, calling it its best reasoning and coding model yet and pricing it at 3.7 Flash's introductory rate: $0.75 per million input tokens and $3.75 per million output tokens — a rate that doubles on January 1, 2027. The same announcement introduced Gemini 3.8 Flash Cyber, a security-tuned variant for autonomous vulnerability discovery and automated patching, available only to vetted defenders through the new Fairwind Program.
This is Google's third Flash release in six weeks, with 3.7 Flash having shipped just three weeks earlier, according to the DeepMind announcement. Both new variants run on the same foundational intelligence, sharpened for coding and reasoning partly through training in cybersecurity, and refined by long-running agentic loops that recursively evaluate the underlying models.
Two variants, one shared core
Gemini 3.8 Flash is positioned as Google's most intelligent workhorse model, with claimed gains over 3.7 Flash across software engineering, agentic tasks and multi-step reasoning in specialized domains, at the same speed as 3.7 Flash. Gemini 3.8 Flash Cyber takes that same core and tunes it for defense: Google describes it as its most capable cybersecurity model, with frontier-level performance in vulnerability detection and automated patching.
| Gemini 3.8 Flash | Gemini 3.8 Flash Cyber | |
|---|---|---|
| Target workload | Coding, agentic tasks, multi-step domain reasoning | Vulnerability discovery and automated patching |
| Who can use it | Gemini API developers, enterprises, AI Pro/Ultra consumers | Trusted defenders only, via the Fairwind Program |
| Price | $0.75 in / $3.75 out per million tokens (introductory) | Not published in the announcement |
| Safety mitigations | Standard safeguards against CBRN and cyber-offense misuse | More permissive cyber mitigations, hence restricted access |
Pricing: $0.75 in, $3.75 out — until January 1, 2027
Gemini 3.8 Flash launches at the same introductory price as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. That rate expires December 31, 2026. Starting January 1, 2027, Google charges $1.50 per million input tokens and $7.50 per million output tokens — a 100% increase on both sides.
| Period | Input per 1M tokens | Output per 1M tokens | Change |
|---|---|---|---|
| Through Dec 31, 2026 | $0.75 | $3.75 | Introductory rate, matches 3.7 Flash |
| From Jan 1, 2027 | $1.50 | $7.50 | +100% input, +100% output |
Google did not publish pricing for 3.8 Flash Cyber or the Fairwind Program. One more caveat: an identical per-token price does not mean an identical per-task cost, because 3.8 Flash can consume more tokens at high effort levels (covered below). For how Gemini's lineup sits against competitors as of July 2026, see our AI model API pricing comparison.
Performance claims: vendor-reported numbers
Google says 3.8 Flash often approaches the performance of higher-cost frontier models. On DeepSWE v1.1, a long-horizon software engineering benchmark, Google reports that 3.8 Flash outperforms most larger frontier models at autonomously solving complex engineering problems end to end, at a fraction of the cost — no score was published. In quantitative and professional domains, Google says it beats 3.7 Flash and other frontier models on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark, again without publishing numbers. The one hard figure is 54.9% on HLE-Verified, the verified subset of the Humanity's Last Exam benchmark, which Google cites as evidence of multi-step reasoning across STEM, humanities and professional fields.
Every one of these results is vendor-reported and unverified, and several omit the comparison models' scores — exactly the pattern we cover in why vendor benchmark scores mislead.
The catch: 3.8 Flash works harder and can burn more tokens
Google attributes the gains to a deliberate design choice — 3.8 Flash works harder. On complex tasks it executes extra reasoning steps and calls tools iteratively, and at times uses more tokens to maximize performance, especially at higher effort levels. If that bites your budget, two levers exist: use lower effort levels to minimize token overhead, or stay on Gemini 3.7 Flash, which Google says remains fully supported for efficiency-first workloads. If you run agents in production, this belongs in your cost model — see our practical guide to working with AI coding agents.
What the announcement does not say
The launch post publishes no context window, rate limits or knowledge cutoff for either variant, and no pricing for Flash Cyber. Anyone migrating on the assumption that 3.8 Flash inherits 3.7 Flash's context limits should verify against the Gemini API documentation before shipping. Google also provides no numeric scores for DeepSWE, Vals Finance Agent V2 or Harvey's Legal Agent Benchmark — only comparative claims.
Gemini 3.8 Flash Cyber and the Fairwind Program
3.8 Flash Cyber is distributed through the Fairwind Program, Google's new channel for giving trusted government authorities, critical infrastructure operators and software maintainers prioritized access. It ships with more permissive cybersecurity mitigations than the standard model, which is why access is restricted. Standard 3.8 Flash retains safeguards against Chemical, Biological, Radiological and Nuclear (CBRN) and cyber-offense misuse, in line with Google's Frontier Safety Framework, and Google reports a significant leap in prompt-injection robustness for the 3.8 family as measured by Gray Swan.
Autonomous vulnerability discovery
On CyberGym, the standard industry benchmark for finding vulnerabilities, Google reports frontier-level autonomous discovery, surpassing both 3.5 Flash Cyber — the previous Cyber model — and significantly larger frontier models. Because CyberGym is limited to C/C++ codebases, Google also ran an internal benchmark spanning complex codebases in 20 programming languages, where 3.8 Flash Cyber reached a success rate exceeding 70%.
Automated patching
On CWE-Bench (named for the Common Weakness Enumeration), a challenging external patching benchmark run by Collinear, 3.8 Flash Cyber achieved a pass@1 of 47.2% — its first submitted patch passing — against 47.8% for a leading frontier model Google does not name, at significantly lower cost. Google places it on the Pareto frontier for that trade-off.
Results inside Google and from partners
- Chrome Security team: 3.8 Flash Cyber produced 2.6 times more correct patches to Chrome vulnerabilities than the best commercial models, which Google notes are much larger.
- Wiz: 7.5–9.7% higher recall on Wiz's internal penetration-testing benchmark, at 2.3–5.2x lower cost compared to other leading frontier models.
- Google Cloud Vulnerability Research: used 3.8 Flash Cyber to find a critical foundational vulnerability in less than 2 hours — research and discovery that Google says usually takes months.
Migrating from 3.7 Flash or 3.6 Flash
- Run your own evals first. The public claims are vendor-reported and partly number-free; your tasks are the benchmark that matters.
- Set effort levels per workload. High effort buys performance and costs tokens; low effort minimizes token overhead.
- Keep 3.7 Flash where efficiency is the constraint. Google explicitly states it remains fully supported for efficiency-first workloads.
- Budget for January 1, 2027. Every token bought at $0.75/$3.75 today costs $1.50/$7.50 in 2027 — build that into annual forecasts now.
- If you are still on 3.6 Flash, note this is the third Flash release in six weeks and the lineup moves fast. Our Gemini 3.6 Flash pricing guide covers where that model landed, and if per-token cost is your only criterion, our DeepSeek V4 Flash coverage tracks a budget alternative.
What to do now
- API developers: start building in Google AI Studio or Android Studio, explore agent-first workflows in Google Antigravity, and generate UIs in Stitch. The DeepMind announcement links the developer docs.
- Enterprises: 3.8 Flash is available in Gemini Enterprise.
- Consumers: AI Pro and Ultra subscribers get 3.8 Flash in the Gemini app, AI Mode in Google Search and Gemini in Google Sheets.
- Security teams: 3.8 Flash Cyber requires applying to the Fairwind Program — the announcement names government authorities, critical infrastructure operators and software maintainers as the intended audience.
Frequently asked questions
How much does Gemini 3.8 Flash cost?
Gemini 3.8 Flash is priced at an introductory $0.75 per million input tokens and $3.75 per million output tokens, matching 3.7 Flash's introductory rate. That price expires December 31, 2026; from January 1, 2027, Google charges $1.50 per million input and $7.50 per million output tokens.
What is the context window of Gemini 3.8 Flash?
Google's September 2, 2026 announcement does not publish a context window, rate limits, or knowledge cutoff for Gemini 3.8 Flash. Developers should verify against the Gemini API documentation before assuming parity with 3.7 Flash.
What is Gemini 3.8 Flash Cyber?
Gemini 3.8 Flash Cyber is a security-tuned variant of 3.8 Flash built for autonomous vulnerability discovery and automated patching. It ships with more permissive cyber-safety mitigations, so it is available only to vetted defenders — government authorities, critical infrastructure operators and software maintainers — through Google's Fairwind Program.
Should I migrate from Gemini 3.7 Flash to 3.8 Flash?
Migrate if you need the coding and multi-step reasoning gains; stay on 3.7 Flash if token cost per task is your main constraint, because 3.8 Flash can consume more tokens at higher effort levels. Google states 3.7 Flash remains fully supported for efficiency-first workloads.