See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite is Google's high-efficiency multimodal model, built for low-latency, high-volume work across text, image, video, audio and PDF inputs. It is the cheap, fast end of the Gemini 3.1 line rather than its reasoning end.

GoogleVerified 24/09/2026Released 07/05/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 231 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.25

field median $0.43

Output / 1M tokens

$1.50

field median $1.80

Context

1,049K

field median 500K

02 / overview

What Gemini 3.1 Flash Lite is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

It is a generally available multimodal model tuned for throughput rather than depth, handling text, images, video, audio and files in the same request. The job it was built for is lightweight agentic work and high-volume pipelines where cost per call matters more than getting the hardest possible answer. It is not the model to reach for when a task needs extended reasoning or careful multi-step judgement.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemini 3.1 Flash Lite first appeared in May 2026 as the efficiency tier of the Gemini 3.1 family, sitting below the fuller Flash and Pro tiers. Pick it when you are running the same operation thousands of times a day and latency is the constraint, classification, extraction, routing, summarisation. Pick something heavier when a single wrong answer is expensive.

How you reach it

The API, the apps it powers, and what its limits let you do.

You reach it through the Gemini API and Google's developer surfaces, calling it like any other Gemini tier. Its context window runs to roughly a million tokens, so you can put long documents, transcripts, whole video files or large PDF sets in front of it in one pass, and its output ceiling is generous enough for long structured responses rather than short replies. The multimodal input range means one endpoint covers audio transcription, document parsing and image understanding without stitching separate services together.

Why it matters

What changes because this exists, or why it does not.

The change here is economic rather than capability-led. Work that was previously batched, cached or skipped because a frontier model made it uneconomic, tagging every support ticket, reading every uploaded PDF, watching every submitted video, becomes something you can run on everything. Nothing new becomes possible, a lot of things become affordable at volume.

Follows Gemini 3.1 Pro Preview Custom Tools. Superseded by Gemini 3.5 Flash. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemini 3.1 Flash Lite API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.25per 1M tokens2026-09-24Check
Output$1.50per 1M tokens2026-09-24Check

It is priced at the budget end of the current multimodal field, with output costing several times input as usual, and the gap between it and frontier tiers is wide enough to change what you are willing to run. At this level the deciding cost is usually volume and context length rather than the rate itself, a million-token prompt is still a million-token prompt.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$12.50
A busy support assistant200M tokens40M tokens$110.00
A document pipeline1000M tokens100M tokens$400.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window1,048,576 tokens
Maximum output65,536 tokens
Modalitiestext, image, video, file, audio
Released07/05/2026
StatusCurrent
Catalogue identifiergoogle/gemini-3.1-flash-lite

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Gemini 2.5 Flash17/06/2025
  2. 02Gemini 2.5 Flash Lite22/07/2025
  3. 03Gemini 3.1 Pro Preview19/02/2026
  4. 04Gemini 3.1 Pro Preview Custom Tools25/02/2026
  5. 05Gemini 3.1 Flash Lite07/05/2026
  6. 06Gemini 3.5 Flash19/05/2026
  7. 07Gemini 3.5 Flash Lite21/07/2026
  8. 08Gemini 3.6 Flash21/07/2026
  9. 09Gemini 3.7 Flash13/08/2026
  10. 10Gemini 3.8 Flash02/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap multimodal inference means AI answer surfaces can afford to read more of your material, more often, including PDFs, product videos and recorded demos rather than just web pages. If your positioning only exists in assets nobody was parsing before, that changes. The practical move is to make sure non-HTML content, documentation PDFs, webinar recordings, product walkthroughs, states plainly what you do and who you do it for, because a model reading at this price point will not work hard to infer it.

Where buyers meet this model

Most buyers meet it indirectly, as the model quietly doing the first pass inside a product they already use rather than as a name they chose. Developers meet it through the Gemini API, where it is the default pick for the high-frequency steps in an agent pipeline. It is also the kind of tier that sits behind fast, cheap AI search and assistant features, where a lighter model handles retrieval and routing before anything heavier is invoked.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

What is Gemini 3.1 Flash Lite used for?
High-volume, low-latency tasks such as classification, extraction, routing and summarisation, plus lightweight agentic steps. It takes text, image, video, audio and PDF input, so it suits pipelines that need to read mixed media cheaply and often.
How does Gemini 3.1 Flash Lite differ from the larger Gemini tiers?
It is the efficiency tier. It trades reasoning depth for speed and cost, which makes it the right pick for repeated operations and the wrong one for tasks where a single wrong answer is expensive.
Can Gemini 3.1 Flash Lite handle long documents and video?
Yes. Its context window runs to roughly a million tokens and it accepts video, audio and file inputs directly, so long transcripts, large PDF sets and recorded media can go in a single request.
Surge45°

Is Gemini 3.1 Flash Lite recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.