See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemini 2.5 Flash

Gemini 2.5 Flash is Google's workhorse model in the Gemini 2.5 family, built for advanced reasoning, coding, mathematics and scientific work, with built-in "thinking" that lets it work through a problem before answering.

GoogleVerified 24/09/2026Released 17/06/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.30

field median $0.30

Output / 1M tokens

$2.50

field median $1.25

Context

1,049K

field median 524K

02 / overview

What Gemini 2.5 Flash is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Gemini 2.5 Flash is a general-purpose reasoning model aimed at the everyday volume work of a product: classification, extraction, code, structured analysis and long-document questions. Google positions it as the default workhorse rather than the top of the range, so it takes on reasoning tasks at scale rather than the hardest single problems where a frontier sibling is the better call.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemini 2.5 Flash first appeared in June 2025 as the mid-tier member of the Gemini 2.5 line, sitting between the cheapest Flash-class options and Google's heavier Pro models. It is the right pick when you are running a lot of requests and still want reasoning behind each one. It is the wrong pick when a single answer carries a lot of weight and you would rather pay for the largest model in the family, or when the job is trivial enough that a lighter model would do.

How you reach it

The API, the apps it powers, and what its limits let you do.

Gemini 2.5 Flash is reached through the Gemini API and Google's developer tooling, and it powers consumer-facing Gemini experiences. It takes text, images, audio, video and files as input, so you can hand it a recording, a screenshot, a video or a long PDF without a separate pipeline. The context window runs to roughly a million tokens, which is enough to hold an entire documentation set, a long transcript or a large codebase in a single request, and output can run long enough for full reports or substantial code.

Why it matters

What changes because this exists, or why it does not.

Gemini 2.5 Flash makes reasoning cheap enough to apply to every request rather than the important ones. Work that used to need a routing layer, a small model for the easy cases and an expensive one for the rest, can now sit on one model. The very large context plus native audio and video input also removes a chunk of preprocessing that teams used to build by hand.

Superseded by Gemini 2.5 Flash Lite. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemini 2.5 Flash API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.30per 1M tokens2026-09-24Check
Output$2.50per 1M tokens2026-09-24Check

Gemini 2.5 Flash is priced as a volume model, cheap enough to sit in the hot path of a product rather than behind a cost gate, with output costing several times more than input as is normal across the field. Against frontier models from the same generation it is an order of magnitude cheaper to run, which is the whole point of it, though genuinely lightweight models still undercut it.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$18.50
A busy support assistant200M tokens40M tokens$160.00
A document pipeline1000M tokens100M tokens$550.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window1,048,576 tokens
Maximum output65,535 tokens
Modalitiesfile, image, text, audio, video
Released17/06/2025
StatusCurrent
Catalogue identifiergoogle/gemini-2.5-flash

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Google

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

The first we track in this line.

  1. 01Gemini 2.5 Flash17/06/2025
  2. 02Gemini 2.5 Flash Lite22/07/2025
  3. 03Gemini 3.5 Flash Lite21/07/2026
  4. 04Gemini 3.6 Flash21/07/2026
  5. 05Gemini 3.7 Flash13/08/2026
  6. 06Gemini 3.8 Flash02/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Because Gemini 2.5 Flash is cheap enough to serve high-volume AI answers, it is likely to be the model actually reading your site and summarising your category for a large share of Google-originated queries. That rewards content that survives compression: clear claims, named products, explicit pricing and comparison detail that a model can lift without inference. The long context also means it can hold your full documentation or a whole competitor set at once, so partial coverage of your own product shows up plainly next to rivals.

Where buyers meet this model

Buyers meet Gemini 2.5 Flash inside the Gemini app and across Google surfaces where Gemini answers appear, often without knowing which model in the family served the response. Developers meet it directly through the Gemini API, where it is a common default for production traffic. Plenty of third-party AI features in other products are also running on it underneath.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Gemini 2.5 Flash good enough for production workloads?
Yes, that is what it was built for. Google describes it as its workhorse model for advanced reasoning, coding, mathematics and scientific tasks, and the pricing is set so you can run it on every request rather than reserving it for hard cases.
What can Gemini 2.5 Flash actually take as input?
Text, images, audio, video and files. That means you can send a recording, a screen capture, a video clip or a long PDF straight to the model without building a separate transcription or extraction step first.
When should I use a Pro model instead?
When a single answer carries real weight and the cost of getting it wrong outweighs the cost of the request. Gemini 2.5 Flash is the volume choice, the heavier models in the family are the choice for the hardest individual problems.
What does the built-in thinking do?
It lets the model work through a problem before it responds, which Google says produces more considered answers. In practice it is what allows a model at this price point to handle reasoning, coding and maths rather than just fast text generation.
Surge45°

Is Gemini 2.5 Flash recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.