See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemini 2.5 Flash Lite

Gemini 2.5 Flash Lite is Google's lightweight reasoning model, the cheapest and fastest tier of the Gemini 2.5 family, built for high-volume work where latency and cost matter more than depth.

GoogleVerified 24/09/2026Released 22/07/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.10

field median $0.30

Output / 1M tokens

$0.40

field median $1.25

Context

1,049K

field median 524K

02 / overview

What Gemini 2.5 Flash Lite is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Gemini 2.5 Flash Lite is a small reasoning model tuned for throughput: faster token generation and lower cost per request than its larger siblings, while still handling text, images, files, audio and video as input. The job it was built for is bulk work, classification, extraction, routing, summarising, tagging, answering the same shape of question a million times. It is not the model to reach for when a task needs sustained, careful reasoning or your best possible answer on a hard problem.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemini 2.5 Flash Lite first appeared in July 2025, sitting at the bottom of the Gemini 2.5 line below Flash and Pro. It is the right pick when you are running a pipeline at scale and the per-request cost is the constraint, or when a user is waiting and the response has to feel immediate. It is the wrong pick for long multi-step reasoning, nuanced writing, or anything where a wrong answer is expensive to catch.

How you reach it

The API, the apps it powers, and what its limits let you do.

Gemini 2.5 Flash Lite is reached through the Gemini API and Google's developer surfaces, and it accepts text, images, files, audio and video in the same request. The context window runs to seven figures of tokens, so you can put an entire document set, a long transcript or hours of media in front of it in one pass rather than chunking, and the output ceiling is generous enough for full structured extractions rather than short replies.

Why it matters

What changes because this exists, or why it does not.

What changes with Gemini 2.5 Flash Lite is the economics of reading everything. Jobs that were previously filtered down before a model ever saw them, every support ticket, every product page, every uploaded file, can now be passed through in full, with a larger model reserved for the cases the small one flags. It is not a capability leap, it is a cost floor that makes whole-corpus processing a routine architectural choice.

Follows Gemini 2.5 Flash. Superseded by Gemini 3.5 Flash Lite. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemini 2.5 Flash Lite API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.10per 1M tokens2026-09-24Check
Output$0.40per 1M tokens2026-09-24Check

Gemini 2.5 Flash Lite is priced at the bottom of the market, among the cheapest models any major provider offers on both input and output, with output costing a small multiple of input. At that level the cost of a single call stops being a design consideration, and the real question becomes how many calls your pipeline makes rather than which model each one uses.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$4.00
A busy support assistant200M tokens40M tokens$36.00
A document pipeline1000M tokens100M tokens$140.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window1,048,576 tokens
Maximum output65,535 tokens
Modalitiestext, image, file, audio, video
Released22/07/2025
StatusCurrent
Catalogue identifiergoogle/gemini-2.5-flash-lite

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Google

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Gemini 2.5 Flash17/06/2025
  2. 02Gemini 2.5 Flash Lite22/07/2025
  3. 03Gemini 3.5 Flash Lite21/07/2026
  4. 04Gemini 3.6 Flash21/07/2026
  5. 05Gemini 3.7 Flash13/08/2026
  6. 06Gemini 3.8 Flash02/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap, fast models get used for the retrieval and filtering stages of AI answers, the part that decides which sources are worth reading closely. That means your pages are increasingly being judged at speed by a model that will not reason its way past ambiguity, so clear claims, plain product naming and facts stated on the page rather than implied matter more than ever. If a small model cannot tell what your product does in one pass, it will not pass you along to the bigger one that writes the answer.

Where buyers meet this model

Buyers rarely meet Gemini 2.5 Flash Lite by name. They meet it as the fast tier behind Google's AI surfaces and as the model quietly doing the first pass inside products built on the Gemini API, the autocomplete, the instant summary, the suggested reply that arrives before you finish reading. If a tool feels immediate and free, a model in this class is usually the reason.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

What is Gemini 2.5 Flash Lite used for?
High-volume, latency-sensitive work: classification, extraction, routing, tagging and summarising at scale, plus the fast first pass in products that need a response to feel instant. It handles text, images, files, audio and video as input.
How does Gemini 2.5 Flash Lite compare with Gemini 2.5 Flash and Pro?
It sits below both. Flash Lite is the cheapest and fastest tier of the Gemini 2.5 family, trading reasoning depth for throughput. Use Flash or Pro when a task needs sustained multi-step reasoning or your best possible answer.
Is Gemini 2.5 Flash Lite cheap enough to run on every request?
For most pipelines, yes. It is priced among the lowest of any major provider's models, which makes whole-corpus processing viable: run everything through Flash Lite, then escalate only the cases it flags to a larger model.
Surge45°

Is Gemini 2.5 Flash Lite recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.