See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemini 3.5 Flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, built to deliver coding and reasoning close to the Pro tier while running at Flash-tier cost and speed. It handles text, images, video, audio and file inputs in a single model.

GoogleVerified 24/09/2026Released 19/05/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 231 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$1.50

field median $0.43

Output / 1M tokens

$9.00

field median $1.80

Context

1,049K

field median 500K

02 / overview

What Gemini 3.5 Flash is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Gemini 3.5 Flash is a general-purpose multimodal model tuned for throughput rather than for the absolute ceiling on hard reasoning. Google optimised it for coding proficiency and for running agents in parallel, so its natural job is the work that happens many times a minute: code generation, extraction, classification, tool calls in a loop. It is not the model you reach for when a single answer needs the strongest reasoning Google can offer, that remains a Pro-tier job.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemini 3.5 Flash first appeared in May 2026 as the efficiency tier of the Gemini 3.5 family, sitting below the Pro models on capability and well below them on cost. It is the right pick when volume, latency or agent concurrency is the constraint, and when near-Pro coding quality is good enough. It is the wrong pick for one-shot work where the answer has to be the best available and cost per call barely matters.

How you reach it

The API, the apps it powers, and what its limits let you do.

Gemini 3.5 Flash is reached through Google's Gemini API and sits behind Google's own consumer and developer surfaces. Its context window runs into the seven figures, which means a full codebase, a long video or a stack of documents can go in as a single prompt rather than being chunked, and its output ceiling is large enough to return substantial generated files in one response. Because it accepts text, image, video, audio and file inputs together, one call can reason across a screen recording, its transcript and the source files behind it.

Why it matters

What changes because this exists, or why it does not.

The practical change is that near-Pro coding quality becomes affordable to run in a loop. Agent designs that were priced out when every step went to a frontier model, parallel branches, retries, self-checking passes, become reasonable at this tier, and long-context work that used to need a retrieval pipeline can often just be pasted in. For teams already on Gemini, it shifts the default: the Pro models become the escalation path rather than the starting point.

Follows Gemini 3.1 Flash Lite. Superseded by Gemini 3.5 Flash Lite. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemini 3.5 Flash API pricingSurge45°
ChargePriceUnitRead onSource
Input$1.50per 1M tokens2026-09-24Check
Output$9.00per 1M tokens2026-09-24Check

Gemini 3.5 Flash is priced as an efficiency model, cheap enough per call that high-volume and multi-step agent workloads stay viable, with output costing several times more than input as is normal across the field. It undercuts frontier Pro-tier pricing by a wide margin while sitting above the cheapest lightweight models, which is the trade you accept for coding quality close to the top tier.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$75.00
A busy support assistant200M tokens40M tokens$660.00
A document pipeline1000M tokens100M tokens$2,400.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window1,048,576 tokens
Maximum output65,536 tokens
Modalitiestext, image, video, file, audio
Released19/05/2026
StatusCurrent
Catalogue identifiergoogle/gemini-3.5-flash

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Gemini 2.5 Flash17/06/2025
  2. 02Gemini 2.5 Flash Lite22/07/2025
  3. 03Gemini 3.1 Pro Preview19/02/2026
  4. 04Gemini 3.1 Pro Preview Custom Tools25/02/2026
  5. 05Gemini 3.1 Flash Lite07/05/2026
  6. 06Gemini 3.5 Flash19/05/2026
  7. 07Gemini 3.5 Flash Lite21/07/2026
  8. 08Gemini 3.6 Flash21/07/2026
  9. 09Gemini 3.7 Flash13/08/2026
  10. 10Gemini 3.8 Flash02/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Efficiency-tier models are the ones actually generating most AI answers at scale, so how your brand is represented here matters more than how it performs against a flagship. Gemini 3.5 Flash reads long documents whole, which rewards clear, complete source pages over fragments tuned for retrieval snippets. If your product documentation, pricing and comparison pages are thin or scattered, a model with this much context will simply fill the gap with whatever third-party page is more complete.

Where buyers meet this model

Buyers meet Gemini 3.5 Flash mostly without naming it: as the model answering inside Google's Gemini app and its AI search surfaces, where speed matters more than a few points of benchmark headroom. Developers meet it directly through the Gemini API, and increasingly inside coding tools and agent frameworks that use it as the default worker model. It is also the tier most SaaS products pick when they embed a Google model in their own product.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Gemini 3.5 Flash good enough for coding, or do I need the Pro tier?
Google positions it as near-Pro on coding, which makes it a reasonable default for most generation, refactoring and review work. Keep a Pro-tier model as the escalation path for the small share of tasks where the hardest reasoning is needed.
What can I do with the large context window?
You can put a whole codebase, a long video, an audio file or a stack of documents into a single prompt instead of chunking and retrieving. That removes a lot of pipeline work for tasks where the model just needs to see everything at once.
What does multimodal mean in practice here?
Gemini 3.5 Flash accepts text, images, video, audio and files, and can reason across them in the same request. In practice that means one call can take a screen recording, its audio and the underlying source files together rather than processing each separately.
Why does an efficiency model matter for AI search visibility?
Fast, cheap models do the bulk of the answering across consumer apps and AI search surfaces, so they shape most of the citations a brand receives. Being well represented in the model that runs at volume is worth more than being well represented in the one that runs occasionally.
Surge45°

Is Gemini 3.5 Flash recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.