See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemma 4 31B

Gemma 4 31B is Google's open-weight multimodal instruct model, a roughly 30.7 billion parameter dense model that reads text, images and video and writes text back. It sits in the Gemma family as the option you run yourself or call cheaply through a host, rather than a hosted frontier model.

GoogleVerified 24/09/2026Released 02/04/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 231 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.09

field median $0.43

Output / 1M tokens

$0.34

field median $1.80

Context

262K

field median 500K

02 / overview

What Gemma 4 31B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Gemma 4 31B is a dense instruction-tuned multimodal model built for general text generation with image and video input alongside it. Google DeepMind ships it with a configurable thinking mode and native function calling, so it is aimed at agent and tool-use work as much as plain chat. It is not a frontier reasoning model and it is not a hosted assistant product, it is a component you deploy into your own stack.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemma 4 31B first appeared in April 2026, positioned in the middle of the Gemma 4 line, above the small models meant for on-device work and below Google's closed Gemini tier. Pick it when you want open weights, multimodal input and function calling without paying hosted frontier rates. It is the wrong pick when a task genuinely needs the strongest available reasoning, or when you need very long generated outputs in a single pass.

How you reach it

The API, the apps it powers, and what its limits let you do.

Gemma 4 31B is reached through the usual API endpoints of whichever provider hosts it, or self-hosted from the open weights. The quarter-million token context comfortably holds a full documentation set, a long transcript or a batch of product pages in one call, and image and video input means you can feed screenshots, interface recordings or product photography without a separate vision step. Output is capped well below the input window, so it suits analysis, extraction and structured responses rather than long-form drafting in one shot.

Why it matters

What changes because this exists, or why it does not.

An open-weight model that takes video and images, holds a very large context and calls functions natively removes a lot of the plumbing teams used to build around smaller open models. Work that previously meant a vision model, a text model and a separate orchestration layer can now sit in one call, at a cost low enough to run across a whole content library rather than a sample of it.

Follows Gemma 3 4B. Superseded by Gemma 4 26B A4B. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemma 4 31B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.09per 1M tokens2026-09-24Check
Output$0.34per 1M tokens2026-09-24Check

Gemma 4 31B is priced at the low end of the market, cheap enough that running it across large document sets or every page on a site is a routine decision rather than a budgeted one. Output costs a few times more than input, as usual, but both sit far below hosted frontier models, and self-hosting the open weights removes per-token cost entirely if you have the infrastructure.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$3.50
A busy support assistant200M tokens40M tokens$31.60
A document pipeline1000M tokens100M tokens$124.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window262,144 tokens
Maximum output16,384 tokens
Modalitiesimage, text, video
Released02/04/2026
StatusCurrent
Catalogue identifiergoogle/gemma-4-31b-it

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Gemma 3 27B12/03/2025
  2. 02Gemma 3 12B13/03/2025
  3. 03Gemma 3 4B13/03/2025
  4. 04Gemma 4 31B02/04/2026
  5. 05Gemma 4 26B A4B03/04/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap open-weight multimodal models mean more companies will run their own retrieval and summarisation over your site rather than relying on a handful of big assistants. That pushes the burden back onto your source material being clean, complete and machine-readable, because there is no single AI engine whose quirks you can optimise for. Assume your pages, docs and screenshots are being read directly, at scale, by systems you will never see.

Where buyers meet this model

Buyers rarely meet Gemma 4 31B by name. They meet it inside products built on top of it, support assistants, content tools and internal search built by teams who wanted open weights and predictable costs, and through the model catalogues of the inference providers that host it.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Gemma 4 31B open weight?
Yes. Gemma 4 31B is released by Google DeepMind as an open-weight instruct model, so you can self-host it or call it through any provider that offers it.
Can Gemma 4 31B handle video?
It accepts image, text and video input and returns text. There is no image or video generation.
What is thinking mode in Gemma 4 31B?
It is a configurable reasoning setting. You can turn extended reasoning on for harder tasks or leave it off when you want faster, cheaper responses.
When should I use something other than Gemma 4 31B?
Choose a larger hosted model when the task needs the strongest available reasoning, or when you need a long single generation, since the maximum output length here is modest relative to the input window.
Surge45°

Is Gemma 4 31B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.