AI Tools Review
Gemini 3.8 Flash launch header graphic from Google's official announcement.

Insights

Gemini 3.8 Flash & Flash Cyber: Benchmarks, Pricing

AI Tools Review Editorial Team3 September 2026Updated 3 September 2026
  • Google
  • Gemini
  • Gemini 3.8 Flash
  • Gemini 3.8 Flash Cyber

Quick answer:

Gemini 3.8 Flash and Gemini 3.8 Flash Cyber launched on 2 September 2026, Google's third Flash-tier update in roughly six weeks. Google calls 3.8 Flash "our best reasoning & coding model yet, at the same speed and low cost of 3.7", and independent cost-efficiency charts back that up: on DeepSWE v1.1, it sits on the Pareto-efficient frontier at roughly 72% accuracy for a fraction of the cost of much larger models. Gemini 3.8 Flash Cyber scores 47.2% Pass@1 on CWE-Bench automated patching - nearly matching the leading frontier model's 47.8% at about a third of the cost per rollout - and reports 2.6x more correct Chrome vulnerability patches than the best commercial models. Pricing is unchanged from 3.7 Flash at $0.75/$3.75 per million tokens (input/output) through 31 December 2026. The Cyber variant is restricted to vetted defenders via Google's Fairwind Program, not generally available.

Three weeks after Gemini 3.7 Flash, Google shipped another Flash-tier update - and this time paired it with a dedicated cybersecurity variant rather than a single general-purpose model. The pairing matters: it lands in the same news cycle as OpenAI's GPT-6 Astra crossing its own Preparedness Framework's Critical cyber threshold, making offensive-versus-defensive cyber capability one of the defining storylines of this model generation.

This review is built from Google's own announcement, its DeepMind evals-methodology pages, and two independently-produced cost-efficiency charts Google published alongside the launch - one built by Datacurve AI on DeepSWE v1.1, the other Google's own CWE-Bench chart for the Cyber variant. Both charts plot accuracy against real dollar cost, which is a more honest signal than a single leaderboard number.

Same-day hands-on coverage of the Gemini 3.8 Flash launch.

Summary

Gemini 3.8 Flash is Google's latest workhorse model for coding and agentic workloads, and Gemini 3.8 Flash Cyber is a parallel release built specifically for vulnerability discovery and patching. Google frames the pairing around one idea: strong reasoning and coding capability, delivered at Flash-tier speed and cost, with a defensive-security specialisation split out into its own restricted-access model rather than bolted onto the general release.

The clearest evidence for the "same speed, low cost" claim isn't a leaderboard score - it's the two cost-efficiency charts Google published. Both show Gemini 3.8 Flash and 3.8 Flash Cyber sitting on or near the Pareto-efficient frontier: matching or beating much larger, far more expensive models at a fraction of the price per task.

  • Best fit: coding agents and terminal-based workflows where cost-per-task matters as much as peak accuracy, and defensive security teams who can access the Cyber variant through Fairwind.
  • Headline chart result: ~72% on DeepSWE v1.1 at a low per-task cost, ahead of Gemini 3.7 Flash's ~65% at a similar price point.
  • Cyber headline: 47.2% Pass@1 on CWE-Bench automated patching, near-matching the leading frontier model's 47.8% at roughly a third of the cost per rollout.
  • Main caveat: Gemini 3.8 Flash Cyber is not a product you can simply sign up for - it is gated behind Fairwind Program vetting for government and critical-infrastructure defenders.

Lineage: A Fast-Moving Flash Line

Gemini 3.7 Flash launched on 13 August 2026, itself only three weeks after Gemini 3.6 Flash. Gemini 3.8 Flash continues that cadence, arriving on 2 September 2026 - Google's third Flash-tier release in around six weeks. That pace suggests Google is iterating primarily through post-training and algorithmic refinement on a stable base rather than shipping a new pretraining run each time, though Google has not published enough architectural detail to confirm this directly.

What is new with this release is the explicit split into two named models at launch: Gemini 3.8 Flash for general use, and Gemini 3.8 Flash Cyber as a dedicated defensive-security variant. Google has shipped cyber-focused evaluations before (the CWE-Bench chart references a "3.5 Flash Cyber" baseline), but pairing a named Cyber release with the mainline Flash launch - rather than treating it as an internal evaluation footnote - is a more deliberate positioning move, and one that lands squarely in the same week as rival frontier labs publishing their own cyber-capability threshold disclosures.

What Shipped: Specs & Availability

Google describes Gemini 3.8 Flash as "our best reasoning & coding model yet, at the same speed and low cost of 3.7" - a deliberately narrow claim about the trade-off curve rather than an assertion of category-leading raw intelligence. The model targets long-horizon software-engineering work, professional and quantitative reasoning, and prompt-injection robustness specifically.

Availability at launch is broad and split by audience. Developers get it in Google Antigravity (Google's agent-first coding environment), Google AI Studio, Android Studio and Stitch. Enterprises get it through the Gemini Enterprise platform. Consumers on Google AI Pro or Ultra subscriptions get it in the Gemini app, Google Search's AI Mode, and Google Sheets. Gemini 3.8 Flash Cyber sits outside that general rollout entirely: access runs through the Fairwind Program, an application-based route aimed at government agencies and critical-infrastructure defenders rather than the standard Gemini API or AI Studio console.

Capabilities Deep Dive

Long-horizon software engineering

Google's own framing is that Gemini 3.8 Flash "outperforms most larger frontier models in autonomously solving complex engineering problems" on DeepSWE v1.1, a long-horizon software-engineering benchmark. The independently-produced Datacurve AI chart (below) supports the specific, narrower version of that claim: Gemini 3.8 Flash does not top the raw-accuracy leaderboard outright - Claude Opus 5, Claude Fable 5 and several others score higher in absolute terms - but it reaches roughly 72% accuracy at a strikingly low average cost per task, putting it on the efficiency frontier alongside models that cost many times more per task to run.

Professional and quantitative reasoning

Google highlights gains on Vals AI's Finance Agent V2 and Harvey's Legal Agent Benchmark, stating 3.8 Flash outperforms both 3.7 Flash and other frontier models on these professional-domain evaluations. On Humanity's Last Exam - Verified, a curated, human-checked subset of the broader HLE benchmark spanning STEM, humanities and professional fields, Gemini 3.8 Flash scores 54.9%, which Google presents as evidence of genuine multi-step reasoning rather than narrow benchmark tuning.

Prompt-injection robustness

Measured against the Gray Swan benchmark - an adversarial-robustness suite specifically designed to probe how models handle malicious instructions embedded in tool outputs or retrieved content - Google reports "significant improvements" for Gemini 3.8 Flash. This matters more than it might first appear: as Flash-tier models get deployed inside increasingly autonomous agent pipelines with access to browsing, code execution and file systems, resistance to indirect prompt injection is one of the more practically important safety properties, distinct from headline reasoning scores.

Benchmarks: Reading the Real Charts

Rather than a single benchmark table, Google published two cost-versus-accuracy charts for this launch - a more informative format than a leaderboard, because it shows where a model sits on the efficiency frontier rather than just its peak score.

Datacurve AI chart plotting DeepSWE V1.1 accuracy against average cost per task on a log scale. Gemini 3.8 Flash sits at roughly 72% accuracy at low cost, on the efficiency frontier above Gemini 3.7 Flash (~65%), Gemini 3.6 Flash and Gemini 3.5 Flash, and near much pricier models including Claude Opus 5, Claude Fable 5, Kimi K3 and Grok 4.6.
DeepSWE V1.1: accuracy vs average cost per task. Gemini 3.8 Flash sits on the Pareto-efficient frontier - matching or beating far pricier frontier models per dollar spent. Source: Datacurve AI (deepswe.datacurve.ai), via Google's Gemini 3.8 Flash launch materials.

Reading the chart in detail: Gemini 3.8 Flash lands at around 72% accuracy for a cost per task well under $1, clearly ahead of Gemini 3.7 Flash's roughly 65% at a broadly comparable cost, and further ahead of Gemini 3.6 Flash (~46%) and Gemini 3.5 Flash (~38%) at even lower prices. Higher up the accuracy axis, models like Claude Opus 5 (~74%) and Claude Fable 5 (~70%) score a few points higher in absolute terms, but at costs several times higher per task - Claude Opus 5 and Claude Fable 5 both sit toward the $10+ end of the cost axis, while Gemini 3.8 Flash sits at the opposite, cheapest end while remaining close to the top of the accuracy scale. That combination - near-frontier accuracy at Flash-tier cost - is the real substance behind Google's "same speed and low cost" framing.

It is worth being precise about what the chart does not show: it is not a claim that Gemini 3.8 Flash is the most capable model available, and Google does not present it as one. Kimi K3, GLM 5.3 and DeepSeek v4 Pro all cluster in a similar cost-and-accuracy region, meaning Gemini 3.8 Flash competes in a genuinely contested part of the efficiency frontier rather than owning it outright.

Gemini 3.8 Flash Cyber in Detail

Gemini 3.8 Flash Cyber is evaluated separately, and Google's own chart for it tells a similarly cost-focused story.

Google chart plotting CWE-Bench Pass@1 automated-patching accuracy against average cost per rollout. Gemini 3.8 Flash Cyber scores 47.2% at roughly $3.60 per rollout, on the Pareto frontier and close to Claude Fable 5's 47.8% at roughly $10.50 per rollout, while beating GPT-5.6 Sol, Gemini 3.7 Flash, Claude Opus 4.8, and open-weight models like Hy4 Preview and DeepSeek V4 Flash on the same cost-accuracy curve.
CWE-Bench automated-patching Pass@1 vs average cost per rollout. Gemini 3.8 Flash Cyber near-matches the leading frontier model's score at roughly a third of the cost. Source: Google DeepMind (deepmind.google/models/evals-methodology/gemini-3-8-flash-cyber).

On CWE-Bench - a benchmark that tests whether a model can automatically patch known vulnerability classes (Common Weakness Enumerations) in real code - Gemini 3.8 Flash Cyber scores a Pass@1 of 47.2% at an average cost of roughly $3.60 per rollout. The named leading frontier model on the same chart, Claude Fable 5, scores 47.8% but costs roughly $10.50 per rollout - nearly three times as much for a fractionally higher score. Gemini 3.7 Flash and GPT-5.6 Sol both cluster around 44% at intermediate cost points, while cheaper open-weight options like Hy4 Preview and DeepSeek V4 Flash trail well behind on accuracy as cost drops toward zero.

Beyond CWE-Bench, Google reports several concrete real-world results for the Cyber variant. On CyberGym, a vulnerability-discovery benchmark, it "surpasses both 3.5 Flash Cyber as well as significantly larger frontier models." In real-world testing, Google reports a success rate exceeding 70% for vulnerability detection across 20 programming languages. Google's Chrome security team reports the model produces 2.6 times more correct patches to Chrome vulnerabilities than the best commercial models it tested against. Cloud-security firm Wiz reports 7.5-9.7 percentage points higher recall for a 2.3-5.2 times lower cost when using the model in its own vulnerability-scanning pipeline - an independent, third-party data point rather than a Google-authored claim, which strengthens its credibility.

A creator's walkthrough of the wider batch of Google AI updates shipped in the same window as Gemini 3.8 Flash.

Safety: CBRN & Cyber Offence Mitigations

Google states that both Gemini 3.8 Flash and Gemini 3.8 Flash Cyber include "safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense" - the same broad domains Google evaluates against its Frontier Safety Framework for every Gemini release. Google has not published a full, itemised Frontier Safety Framework capability-level breakdown for this specific pair of models in the same detail as its Gemini 3.7 Flash model card, which is a real disclosure gap worth flagging rather than assuming parity.

For the Cyber variant specifically, Google describes using "more permissive mitigations for cybersecurity" - a deliberate trade-off that allows the model to perform genuinely useful offensive-adjacent tasks (vulnerability discovery, exploit-pattern recognition needed to write an effective patch) that would otherwise be blocked by the safeguards applied to the general-release Gemini 3.8 Flash. Google's chosen control for that trade-off is access restriction rather than capability restriction: the Fairwind Program vets who can use the more permissive model at all, rather than trying to build a classifier that distinguishes defensive from offensive intent at the point of use. That is a coherent approach, but it does concentrate risk on the vetting process itself - Google has not published details of what that vetting involves beyond describing it as "application-based access for government and critical infrastructure."

Real-World Reception

Launch-day creator coverage focused heavily on the "Google is back" framing - a reference to Google closing what many reviewers had characterised as a capability gap against OpenAI and Anthropic's most recent releases. That framing is directionally supported by the DeepSWE cost-efficiency chart, but it is worth treating with the same scepticism as any vendor-adjacent headline: the chart shows strong cost-efficiency, not an outright accuracy lead over every rival model, and several open- and closed-weight competitors sit in a genuinely similar region of the frontier.

The more durable real-world signal is Wiz's independent recall and cost data for the Cyber variant, cited above, because it comes from a security vendor actually running the model in a production-adjacent vulnerability-scanning workflow rather than a benchmark harness. That kind of third-party operational data point is rarer and more valuable than another leaderboard screenshot, and it is the strongest single piece of evidence in this launch that the efficiency-frontier framing holds up outside Google's own charts.

Pricing

Model and periodInput $/1MOutput $/1M
Gemini 3.8 Flash, introductory rate through 31 Dec 2026$0.75$3.75
Gemini 3.8 Flash, standard rate from 1 Jan 2027$1.50$7.50
Gemini 3.8 Flash CyberNot publicly priced - Fairwind Program access only

Gemini 3.8 Flash carries exactly the same introductory pricing as Gemini 3.7 Flash did at its own launch - Google is not using price as the differentiator between the two Flash generations, which reinforces that this release is positioned as a capability and efficiency refinement on a stable cost base rather than a repriced product. As with 3.7 Flash, the introductory rate is scheduled to double on 1 January 2027, so teams budgeting past that date should plan around the standard $1.50/$7.50 rate.

Limitations

  • Not an outright accuracy leader: on DeepSWE v1.1, several larger models (Claude Opus 5, Claude Fable 5, Kimi K3) score higher in absolute accuracy - Gemini 3.8 Flash's advantage is cost-efficiency, not peak capability.
  • Cyber variant is not generally accessible: Gemini 3.8 Flash Cyber requires Fairwind Program vetting; most developers cannot simply call it via the standard API.
  • Limited public safety disclosure: Google has not published a full itemised Frontier Safety Framework capability-level breakdown for this launch comparable to earlier Gemini model cards.
  • The price is promotional: the $0.75/$3.75 rates are scheduled to double on 1 January 2027.
  • No published pricing for the Cyber variant, making direct cost comparisons against competing defensive-security tools difficult for prospective Fairwind applicants.

How It Compares

Against its immediate predecessor Gemini 3.7 Flash, the improvement is clear and chart-verified on DeepSWE v1.1 (roughly 72% vs 65% at a similar cost point) at unchanged introductory pricing. Against Claude Fable 5.1 and Claude Opus 5, Gemini 3.8 Flash trails on raw accuracy but wins decisively on cost-per-task - the right comparison depends entirely on whether a workload is accuracy-bound or budget-bound. On the cyber-capability axis, the contrast with OpenAI's GPT-6 Astra is one of framing as much as capability: Astra's system card discloses crossing OpenAI's own Critical cyber threshold under tight, layered safeguards, while Google positions Gemini 3.8 Flash Cyber explicitly as a defensive tool distributed only to vetted defenders - two different labs making two different but overlapping bets on how to handle rapidly advancing offensive-adjacent cyber capability.

The more useful framing than "which model wins" is, again, cost-per-task: Gemini 3.8 Flash's specific strength is landing near the top of several efficiency frontiers rather than owning the absolute top of any single accuracy leaderboard, which lines up with Google's own "same speed and low cost" framing for the launch.

Who Should Use It

Choose Gemini 3.8 Flash if you are running coding agents or terminal-based workflows at scale where cost-per-task compounds quickly, and where near-frontier accuracy at a fraction of the price beats squeezing out the last few points of peak accuracy - particularly if you are already inside Google's ecosystem via Antigravity, AI Studio or Gemini Enterprise.

Apply for Fairwind access if you run vulnerability-management or patching workflows for a government agency or critical-infrastructure operator and can use a genuinely more capable, more permissively-configured model under vetted access. Look elsewhere if you need the single highest accuracy score available regardless of cost, or you need a defensive cyber tool today without going through an application process - Gemini 3.8 Flash Cyber is not a self-serve product.

The Bottom Line

Gemini 3.8 Flash and 3.8 Flash Cyber are best understood through the two cost-efficiency charts Google chose to lead with rather than a single headline benchmark score. On DeepSWE v1.1 and CWE-Bench alike, the pattern repeats: near-frontier accuracy at a meaningfully lower cost than the models that narrowly out-score it. That is a genuinely useful, verifiable claim, and the Wiz recall/cost data point for the Cyber variant is the kind of independent evidence that makes it more credible than launch-day marketing alone.

What is missing is a full public safety disclosure comparable to Google's own Gemini 3.7 Flash model card, and any public pricing for the Cyber variant, which remains gated behind Fairwind Program vetting rather than generally available. Choose Gemini 3.8 Flash when cost-per-task efficiency matters more than owning the top spot on a leaderboard, and treat the Cyber variant's restricted access as a deliberate, disclosed trade-off rather than an oversight.

Last updated: 3 September 2026. Sources: Google's official announcement, the DeepSWE v1.1 cost-efficiency chart (Datacurve AI, deepswe.datacurve.ai) and Google's CWE-Bench evaluation methodology page (deepmind.google/models/evals-methodology/gemini-3-8-flash-cyber).

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

When did Gemini 3.8 Flash launch, and how is it different from Gemini 3.8 Flash Cyber?
Google released both on 2 September 2026, three weeks after Gemini 3.7 Flash - the third Flash-tier update in about six weeks. Gemini 3.8 Flash is the general-purpose model, positioned by Google as its best reasoning and coding model yet at the same speed and low cost as 3.7 Flash. Gemini 3.8 Flash Cyber is a specialised variant tuned for vulnerability discovery and automated patching, deployed with more permissive cybersecurity mitigations and restricted to vetted defenders through Google's Fairwind Program rather than general release.
How much does Gemini 3.8 Flash cost?
Introductory pricing through 31 December 2026 is $0.75 per million input tokens and $3.75 per million output tokens - unchanged from Gemini 3.7 Flash's introductory rate. Standard pricing rises to $1.50 input and $7.50 output per million tokens from 1 January 2027. Google has not published separate public pricing for Gemini 3.8 Flash Cyber, which is distributed through the application-based Fairwind Program rather than general API access.
What do Google's own benchmark charts actually show?
On the DeepSWE v1.1 long-horizon coding benchmark, Datacurve AI's cost-efficiency chart places Gemini 3.8 Flash at roughly 72% accuracy at a very low average cost per task, ahead of Gemini 3.7 Flash (around 65% at a similar cost) and sitting on the Pareto-efficient frontier alongside far pricier models like Claude Opus 5 and Claude Fable 5. On CWE-Bench automated patching, Gemini 3.8 Flash Cyber scores a Pass@1 of 47.2%, essentially matching the leading frontier model's 47.8% (Claude Fable 5, at roughly $10.50 per rollout) while costing only about $3.60 per rollout - a large cost advantage for a near-identical patching accuracy.
Is Gemini 3.8 Flash Cyber a dangerous offensive cyber tool?
Google frames it as a defensive tool with capability safeguards rather than an offensive one. Both models carry safeguards against misuse in Chemical, Biological, Radiological and Nuclear (CBRN) domains and cyber offence. The Cyber variant specifically uses more permissive mitigations to allow legitimate vulnerability discovery and patching at scale, but access is restricted to trusted defenders vetted through the Fairwind Program - it is not available through the standard Gemini API. Google reports real defensive wins from this restricted access, including 2.6 times more correct Chrome vulnerability patches than the best commercial models and a 70%+ real-world vulnerability detection success rate across 20 programming languages.
Where can I use Gemini 3.8 Flash?
Gemini 3.8 Flash is available to developers through Google Antigravity, Google AI Studio, Android Studio and Stitch; to enterprises through Gemini Enterprise; and to consumers through the Gemini app, Google Search's AI Mode and Google Sheets for Google AI Pro and Ultra subscribers. Gemini 3.8 Flash Cyber is available only through the Fairwind Program, an application-based access route aimed at government and critical-infrastructure defenders.
AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.