AI Tools Review

Gemini 3.5 Flash

By Google

Released: 2026-05-19

LLM
Multimodal
Agents
Google
Speed
Paid
New

Gemini 3.5 Flash is Google's fast multimodal model and - with Gemini 3.5 Pro repeatedly delayed - the only shipped model of the 3.5 generation. It leads frontier rivals on multimodal and MCP tool-use benchmarks, claims ~4x faster output than other frontier models, and costs $1.50/$9 per million tokens.

Visit Gemini 3.5 Flash

Built for speed

Google claims roughly 4× faster output tokens per second than other frontier models, and Flash's whole design brief follows from that: a model quick enough to sit inside agent loops, multimodal pipelines and user-facing products without making anyone wait.

Multimodal and agentic leader

On Google's launch table Flash tops the row on MCP Atlas (83.6%), Toolathlon (56.5%), Finance Agent v2 (57.9%), CharXiv Reasoning (84.2%) and MMMU-Pro (83.6%), an unusual spread of wins for a Flash-tier model.

Value over peak power

At $1.50 per million input tokens and $9 per million output, with an Artificial Analysis cost per Intelligence Index task of $0.59, Flash scores 85 on our Value For Money Index. It trails Opus 4.7 and GPT-5.5 on the hardest work, but rarely on price.

Gemini 3.5 Flash arrived at Google I/O on 19 May 2026 under the banner of "frontier intelligence with action", and, nearly two months on, it remains the only shipped model of Google's 3.5 generation, with Gemini 3.5 Pro repeatedly delayed. That accident of scheduling makes Flash unusually important: it is currently the whole 3.5 story. This review looks at what it does well (speed, multimodal reasoning, agentic tool use) where it trails the frontier, and whether it earns its 85 Value For Money score.

What Gemini 3.5 Flash is, and why it's carrying the whole generation

Gemini 3.5 Flash benchmark results summary card: positioning and best-for list in the campaign index-card style.
At a glance: where this model fits.

Gemini 3.5 Flash was announced and shipped on the same day, 19 May 2026, on the Google I/O stage, positioned under the tagline "frontier intelligence with action". In Google's naming scheme, Flash is the fast, cheaper tier of a model generation; the flagship reasoning model, Gemini 3.5 Pro, was promised for "next month". That promise slipped. June came and went, then early July, and Pro is now reported to be targeting 17 July 2026 after what has been described as a base-model rebuild. As of today, Flash is the only Gemini 3.5 model anyone can actually use.

That context matters for how you read this review. Flash was never meant to be Google's best model: it was meant to be the quick one, the workhorse that handles high-volume traffic while Pro handles the hard problems. But with Pro absent, Flash has spent almost two months being judged against full-fat frontier models from Anthropic and OpenAI, a comparison Google itself invited by publishing a launch benchmark table that lines Flash up against Claude Opus 4.7 and GPT-5.5.

The surprising thing is how well it survives that comparison in places. On multimodal understanding and agentic tool use, Flash doesn't just hold its own against heavier models: it leads several rows outright. On the hardest coding and abstract reasoning benchmarks it clearly trails, which is exactly what you'd expect from a speed-first model. The result is a genuinely interesting product: not a cut-down version of something better, but a model with a distinct shape (fast, perceptive, action-oriented) that happens to be standing in for a flagship that hasn't arrived. It ships under Google's Frontier Safety Framework, with strengthened cyber and CBRN safeguards, consistent with how Google has framed the whole 3.5 family.

Speed and multimodal strengths

Speed is the headline claim. Google says Gemini 3.5 Flash produces output tokens roughly four times faster than other frontier models. We treat vendor speed claims with caution as a rule, but the direction of the claim fits the product: Flash models have always been tuned for throughput, and this one is explicitly pitched at workloads where latency is a feature, not a nicety: chat interfaces, agent loops that make many model calls per task, document pipelines processing thousands of items, and anything user-facing where a spinning cursor costs you engagement.

Speed alone would make Flash a utility model. What makes it more interesting is that the speed comes attached to genuinely strong multimodal reasoning. On CharXiv Reasoning, which tests understanding of charts and figures from scientific papers, Flash scores 84.2%, the best result in Google's comparison table, ahead of Claude Opus 4.7 and GPT-5.5. On MMMU-Pro, a broad multimodal understanding benchmark, it posts 83.6%, again leading the row. These are not soft benchmarks, and they are not the sort of rows a fast-tier model normally wins.

The practical upshot is that Flash is arguably at its best when you hand it something visual and ask it to reason, quickly, at scale: extracting structure from charts, reading dense documents with figures, interpreting screenshots inside an agent workflow. Its 78.4% on OSWorld-Verified, a benchmark of computer-use tasks in real operating-system environments, reinforces the same picture: this is a model built to perceive an environment and act in it. "Frontier intelligence with action" is marketing language, but on the multimodal-plus-agentic axis, the benchmark evidence broadly backs it up.

  • Google claims ~4× faster output tokens/sec than other frontier models
  • CharXiv Reasoning 84.2% and MMMU-Pro 83.6%, both lead Google's comparison table
  • OSWorld-Verified 78.4% on real computer-use tasks

Benchmarks: leads on multimodal and MCP, trails the frontier on hard coding

Google's launch table compares Flash against Gemini 3 Flash, Gemini 3.1 Pro, Claude Sonnet 4.6, Claude Opus 4.7 and GPT-5.5, and the pattern is unusually clean. Where the task involves tools, environments or images, Flash leads. Where the task is pure, hard coding or expert-level reasoning, the heavyweight models pull ahead.

Start with the wins. MCP Atlas, which measures how well a model drives real tools through the Model Context Protocol, comes in at 83.6%, top of the row. Toolathlon, a long-horizon tool-use benchmark, is 56.5%, also leading. Finance Agent v2, at 57.9%, leads again. Add the multimodal wins covered above and Flash tops five rows of Google's own table against models that cost considerably more to run. For anyone building agents that orchestrate MCP servers or external tools, these are the most relevant numbers on the page.

Now the losses. On SWE-Bench Pro (public, single attempt), which tests resolving real software engineering issues, Flash scores 55.1% against Claude Opus 4.7's 64.3%, a nine-point gap that will be felt on genuinely hard coding work. Terminal-Bench 2.1 is closer: 76.2% versus GPT-5.5's 78.2%. On GDPval-AA, an Elo-scored measure of economically valuable knowledge work, Flash's 1656 sits well below GPT-5.5's 1769. Blueprint-Bench 2, at 33.6%, is low in absolute terms. And on the general-intelligence composites, Flash posts a respectable 40.2% on Humanity's Last Exam (without tools) and a notable 72.1% on ARC-AGI-2.

Artificial Analysis's independent Intelligence Index v4.1 places Flash at 50, one point below GPT-5.6 Luna and GLM-5.2, both at 51. That single point understates the gap on the hardest tasks but captures something true: on aggregate, Flash is within touching distance of models positioned well above it, at a fraction of the running cost.

Where it sits on the Intelligence Index

Artificial Analysis Intelligence Index v4.1 across every scored frontier model: this model highlighted.

Source: Artificial Analysis (9 July 2026). Interactive, hover any bar. Explore the full benchmarks →

The long-context story: a big window with a recall caveat

Gemini 3.5 Flash ships with a context window of roughly one million tokens, 1,048,576 as observed on OpenRouter's serving endpoints. On paper that is enough to hold entire codebases, lengthy contract sets or hours of transcripts in a single request, and long context has been a Gemini selling point for generations. The window is real. The question, as always, is how much of it the model can actually use.

Google's own launch table gives an unusually honest answer. On MRCR v2 (a multi-round co-reference benchmark that tests whether a model can retrieve and use information buried deep in its context), Flash scores 77.3% at 128k tokens. That is a strong result: at 128k, which already covers a small book or a substantial repository, recall is dependable. But at the full 1M-token pointwise setting, the score falls to 26.6%. That is a steep drop, and it is the single most important caveat in this review.

The practical rule that falls out of these numbers: treat Flash as an excellent 128k-context model with an emergency reserve, not as a true million-token reasoner. If your workload genuinely depends on precise recall across hundreds of thousands of tokens (needle-in-haystack retrieval from massive document sets, say), you should test at your actual context lengths before committing, and you may be better served by retrieval pipelines that keep the working context under 128k. The window size is a capability ceiling, not a performance guarantee. Window size and usable recall are different things, and Flash's own scorecard says so.

  • ~1M-token context window (1,048,576 observed on OpenRouter)
  • MRCR v2 at 128k: 77.3%, strong, dependable recall
  • MRCR v2 at 1M pointwise: 26.6%, a steep drop; test before relying on the full window

Pricing and value

Gemini 3.5 Flash benchmark results specification card: intelligence, coding index, cost per task, API pricing, context window and value score.
The numbers in one card: data from our benchmarks tracker.

Flash's API pricing is $1.50 per million input tokens and $9 per million output tokens. Those are not the cheapest numbers in the market, but they buy a model that leads frontier-class competition on five benchmark rows, and the value calculation is where Flash's case becomes hard to argue with.

Artificial Analysis puts Flash's cost per Intelligence Index task at $0.59. Combine that with its Intelligence Index score of 50, one point behind GPT-5.6 Luna and GLM-5.2, and you get a model delivering near-parity aggregate intelligence at a running cost that changes what is economically feasible. Workloads that would be marginal or loss-making on a flagship model (high-volume document processing, always-on agents, per-user personalisation, bulk multimodal analysis) become straightforwardly viable.

On our own Value For Money Index, Flash scores 85, one of the stronger results we have recorded. The reasoning is simple: value is intelligence per pound, adjusted for what the intelligence is good at, and Flash's strengths (multimodal, tool use, speed) map onto the highest-volume categories of real-world API spend. Most production tokens are not spent solving Opus-grade engineering problems; they are spent reading documents, driving tools, answering users and processing images. On that mix, Flash's price-performance is excellent.

The honest counterweight: output tokens at $9 per million are meaningfully dearer than input, so verbose agentic workloads, long chains of tool calls with lengthy reasoning between them, will cost more than the input price suggests. And if your workload lives at the hard end of coding and reasoning, the money saved on Flash may be spent again on retries and human review. Value depends on fit; for the right workloads, Flash's fit is unusually good.

Cost per Intelligence Index task, in context

What a unit of benchmarked work actually costs across the field. Lower is better.

Source: Artificial Analysis (9 July 2026). Interactive, hover any bar. Explore the full benchmarks →

Limitations, and the Gemini 3.5 Pro shadow

Flash's limitations are legible from its scorecard. Hard software engineering is the clearest: 55.1% on SWE-Bench Pro against Opus 4.7's 64.3% means that on genuinely difficult, real-world coding tasks, roughly one task in ten that Opus resolves, Flash won't. GDPval-AA tells a similar story for expert knowledge work: 1656 Elo versus GPT-5.5's 1769 is not a rounding error. Blueprint-Bench 2 at 33.6% is weak in absolute terms. And the long-context recall drop, 77.3% at 128k collapsing to 26.6% at 1M, means the headline window should be treated with care, as covered above.

None of this is disqualifying for a Flash-tier model. What complicates the picture is the model that isn't here. Gemini 3.5 Pro was promised within a month of I/O, slipped to June, then July, and is now reported to be targeting 17 July 2026 following a base-model rebuild. That rebuild detail is worth pausing on: rebuilding a base model mid-cycle is not a scheduling adjustment, it suggests Google decided the original wasn't good enough to ship. Whatever the cause, the effect is that every weakness in Flash's profile currently comes with an implicit IOU attached.

That creates a genuine decision problem for buyers. If Pro lands in the coming days and addresses the hard-coding and reasoning gaps, teams that committed heavily to Opus 4.7 or GPT-5.5 for those workloads may find themselves re-evaluating within weeks. If Pro slips again, or ships underwhelming after its rebuild, Flash's gaps become the 3.5 generation's gaps for longer. Our advice is not to defer decisions indefinitely, Flash is judgeable on its own merits today, and its strengths are real, but to keep the frontier-tier decision loosely held until Pro actually ships and can be measured. We will review Gemini 3.5 Pro separately when it does.

Verdict

Gemini 3.5 Flash benchmark results verdict card with our one-line assessment.
The verdict, briefly.

Gemini 3.5 Flash is the best version yet of a familiar idea: a fast, inexpensive model that is good enough for most work and exceptional at a few things. What is new in this generation is where the exceptional bits sit. Leading frontier-class competition on MCP Atlas, Toolathlon, Finance Agent v2, CharXiv Reasoning and MMMU-Pro is not normal for a speed-tier model, and it makes Flash a first-choice candidate, not a budget compromise, for agentic and multimodal production systems.

The weaknesses are equally clear and should be planned around rather than discovered. Hard software engineering belongs on Opus 4.7 or GPT-5.5 for now. Expert-grade knowledge work shows a real Elo gap. The million-token context window should be treated as a 128k window with headroom, because that is what the recall numbers support. And the whole assessment sits in the shadow of Gemini 3.5 Pro, reportedly days away after a slip from "next month" to mid-July and a base-model rebuild.

Our bottom line: at $1.50 in and $9 out, with an Artificial Analysis Intelligence Index of 50 at $0.59 per task and a Value For Money score of 85, Flash is one of the most economically sensible models on the market for high-volume, tool-driven, multimodal work. If that describes your workload, adopt it now and don't wait for Pro. If your workload lives at the frontier of coding and reasoning, use something heavier today, and watch 17 July, because the 3.5 generation's real flagship argument hasn't been made yet.

Gemini 3.5 Flash benchmark results

Artificial Analysis Intelligence Index v4.150

One point below GPT-5.6 Luna and GLM-5.2 (51); cost per task $0.59

MCP Atlas83.6%

Leads Google's comparison table

CharXiv Reasoning84.2%

Leads, chart and figure reasoning

MMMU-Pro83.6%

Leads, multimodal understanding

Terminal-Bench 2.176.2%

GPT-5.5 leads at 78.2%

SWE-bench Pro (public, single attempt)55.1%

Claude Opus 4.7 leads at 64.3%

Google official launch table plus Artificial Analysis independent testing, May–July 2026. Comparison models: Gemini 3 Flash, Gemini 3.1 Pro, Claude Sonnet 4.6, Claude Opus 4.7, GPT-5.5.

Where Gemini 3.5 Flash fits

Agentic tool orchestration over MCP

Flash's table-leading 83.6% on MCP Atlas and 56.5% on Toolathlon, combined with its ~4× output speed claim, make it a natural engine for agents that chain many tool calls, where per-step latency and per-token cost compound across every task.

High-volume document and chart analysis

With 84.2% on CharXiv Reasoning and 83.6% on MMMU-Pro, Flash excels at extracting structure from figures, charts and dense documents. At $1.50 per million input tokens, bulk pipelines that would be uneconomic on flagship models become routine.

Computer-use and workflow automation

A 78.4% score on OSWorld-Verified supports agents that perceive screens and act in real operating-system environments, form-filling, testing flows and desktop automation where speed keeps loops responsive.

Financial and operational agents

Flash leads its comparison row on Finance Agent v2 at 57.9%, suggesting genuine competence at multi-step financial workflows (data gathering, reconciliation-style tasks and tool-driven analysis) at a cost that suits always-on deployment.

Latency-sensitive user-facing products

Chat assistants, in-app copilots and real-time features benefit directly from Flash's speed-first design. An Intelligence Index of 50 means responsiveness comes without the steep capability sacrifice earlier fast-tier models demanded.

Sources & further reading

Google Model Timeline

Google Omni
Google Spark
Gemini 3.5 FlashCurrent

1M tokens context

Gemini 3.5 FlashCurrent
Google: Gemini 3 Flash Preview

1,049k tokens context

Google: Gemini 3 Pro Preview

1,049k tokens context

Google: Gemini 2.5 Flash Preview 09-2025

1,049k tokens context

Google: Gemini 2.5 Flash Lite Preview 09-2025

1,049k tokens context

Google: Gemini 2.5 Flash Lite

1,049k tokens context

Google: Gemma 3n 2B (free)

8k tokens context

Google: Gemini 2.5 Flash

1,049k tokens context

Google: Gemini 2.5 Pro

1,049k tokens context

Google: Gemini 2.5 Pro Preview 06-05

1,049k tokens context

Google: Gemma 3n 4B (free)

8k tokens context

Google: Gemma 3n 4B

33k tokens context

Google: Gemini 2.5 Pro Preview 05-06

1,049k tokens context

Google: Gemma 3 4B (free)

33k tokens context

Google: Gemma 3 4B

96k tokens context

Google: Gemma 3 12B (free)

33k tokens context

Google: Gemma 3 12B

131k tokens context

Google: Gemma 3 27B (free)

131k tokens context

Google: Gemma 3 27B

96k tokens context

Google: Gemini 2.0 Flash Lite

1,049k tokens context

Google: Gemini 2.0 Flash

1,049k tokens context

Google: Gemini 2.0 Flash Experimental (free)

1,049k tokens context

Google: Gemma 2 27B

8k tokens context

Google: Gemma 2 9B

8k tokens context

Frequently Asked Questions

Is Gemini 3.5 Flash actually a frontier model, or a budget tier?

Both, depending on the axis. On multimodal reasoning and agentic tool use it leads Google's launch table against Claude Opus 4.7 and GPT-5.5: CharXiv Reasoning 84.2%, MMMU-Pro 83.6%, MCP Atlas 83.6%. On the hardest coding and knowledge work it clearly trails: 55.1% on SWE-Bench Pro versus Opus 4.7's 64.3%, and 1656 Elo on GDPval-AA versus GPT-5.5's 1769. It is frontier-class in its lanes, fast tier outside them.

When is Gemini 3.5 Pro coming, and should I wait for it?

Pro was promised a month after the 19 May 2026 I/O launch, slipped to June and then July, and is now reported to be targeting 17 July after a base-model rebuild. If your workload matches Flash's strengths (agents, multimodal, high volume) adopt Flash now; Pro won't change that calculus. If you need frontier coding or reasoning, use Opus 4.7 or GPT-5.5 today and re-evaluate once Pro ships and can be independently measured.

Can I really use the 1M-token context window?

With care. The window is real, 1,048,576 tokens observed on OpenRouter endpoints, but recall degrades sharply with depth. Flash scores a strong 77.3% on MRCR v2 at 128k tokens, falling to 26.6% at the full 1M pointwise setting. Treat it as an excellent 128k model with headroom, and test at your actual context lengths before relying on deep recall.

How much does Gemini 3.5 Flash cost to run?

API pricing is $1.50 per million input tokens and $9 per million output tokens. Artificial Analysis measures its cost per Intelligence Index task at $0.59, against an Index score of 50, one point below GPT-5.6 Luna and GLM-5.2. On our Value For Money Index it scores 85. Note that output tokens cost six times input, so verbose agentic workloads cost more than the input price suggests.

Is Gemini 3.5 Flash good enough for software engineering?

For routine coding, plausibly: its 76.2% on Terminal-Bench 2.1 sits only two points behind GPT-5.5. For hard, real-world engineering the gap widens: 55.1% on SWE-Bench Pro (public, single attempt) against Claude Opus 4.7's 64.3%. A sensible pattern is Flash for high-volume, tool-driven development tasks and a heavier model for the problems where a failed attempt is expensive.

Specifications

pricing$1.50/$9 per 1M tokens
context Window1M tokens

AI Evaluation

4.6
Expert Rating
Text4.6/5
Coding4.2/5

Frontier-class in its lanes - multimodal, tool orchestration, speed - and honest mid-tier outside them. One of the most economically sensible models for high-volume agentic and visual work.

Pros

  • Leads MCP Atlas, CharXiv, MMMU-Pro rows
  • ~4x output speed claim
  • Strong value at $0.59 per measured task

Cons

  • Trails badly on hardest coding (SWE Pro 55.1%)
  • 1M window but weak recall at full depth (26.6% MRCR)
  • Overshadowed by the delayed 3.5 Pro