See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemma 3 4B

Gemma 3 4B is a small open-weight model from Google that takes both text and images as input and returns text. It is the compact member of the Gemma 3 family, built to run cheaply at scale rather than to win on hard reasoning.

GoogleVerified 24/09/2026Released 13/03/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 179 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.05

field median $0.46

Output / 1M tokens

$0.10

field median $1.82

Context

131K

field median 524K

02 / overview

What Gemma 3 4B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Gemma 3 4B is a small multimodal language model: vision-language in, text out, with support for over 140 languages and improved maths, reasoning and chat behaviour over earlier Gemma releases. It is built for high-volume, well-defined work such as classification, extraction, tagging, summarising and routing. It is not the model to reach for when a task needs deep multi-step reasoning or long-form authored output.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemma 3 4B first appeared in March 2025 as part of the Gemma 3 launch, sitting at the small end of the family alongside larger Gemma 3 sizes. Pick it when the job is repetitive, the volume is high and the cost per call matters more than the last few points of quality. Pick a larger sibling when the task involves genuine reasoning or the output is going in front of a customer unedited.

How you reach it

The API, the apps it powers, and what its limits let you do.

Gemma 3 4B is available as open weights, so it can be self-hosted, and it is also served through hosted inference providers with a standard token-priced API. The context window is long enough to hold a full document set, a long support thread or a batch of records in a single call, and image input means screenshots, scanned pages and product photos can go in alongside the text. Output length is capped well below the input budget, which suits structured returns rather than long essays.

Why it matters

What changes because this exists, or why it does not.

Gemma 3 4B makes it reasonable to run a language model over every record rather than a sample. Work that used to be awkward, tagging an entire content archive, reading images across a whole product catalogue, screening thousands of inbound messages before a bigger model sees them, becomes a matter of throughput rather than budget. As a single-call model it is unremarkable, its value is in the volume it absorbs.

Follows Gemma 3 12B, and is the newest in its line. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemma 3 4B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.05per 1M tokens2026-09-24Check
Output$0.10per 1M tokens2026-09-24Check

Gemma 3 4B sits at the cheap end of the market, priced for bulk work where a call happens thousands of times a day rather than once. Output costs twice what input does, which is the usual shape, and self-hosting the open weights removes per-token cost entirely if you have the infrastructure.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$1.50
A busy support assistant200M tokens40M tokens$14.00
A document pipeline1000M tokens100M tokens$60.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output16,384 tokens
Modalitiestext, image
Released13/03/2025
StatusCurrent
Catalogue identifiergoogle/gemma-3-4b-it

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

← Gemma 3 12B

Came after

Nothing newer in this line yet.

  1. 01Gemma 3 12B13/03/2025
  2. 02Gemma 3 4B13/03/2025

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Gemma 3 4B does not answer buyer questions, so it will not cite your brand directly. What it changes is the pipeline behind the answers: models this cheap mean more of the web gets read, parsed and classified in bulk, including your documentation, pricing pages and support content. Structure that content so a small model can extract facts from it without inference, because increasingly a small model is what reads it first.

Where buyers meet this model

Buyers rarely meet Gemma 3 4B as a named product. They meet it inside software that embeds it, as the layer that classifies a ticket, extracts fields from an uploaded document or writes an image caption, and engineers meet it directly through the open weights or a hosted inference API. It is not the model behind a consumer chat assistant or an AI search surface.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Can Gemma 3 4B read images?
Yes. Gemma 3 accepts vision-language input and returns text, so screenshots, scanned documents and photographs can be passed in alongside a prompt. It does not generate images.
Is Gemma 3 4B good enough for customer-facing output?
Treat it as a worker model rather than an author. It is well suited to classification, extraction and summarising behind the scenes, and anything going in front of a customer unedited is better handled by a larger model in the family.
How many languages does Gemma 3 4B handle?
Google states support for over 140 languages, which makes it a practical option for multilingual tagging, triage and extraction work across a large corpus.
Surge45°

Is Gemma 3 4B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.