See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Google model°

Gemma 4 26B A4B

Gemma 4 26B A4B is an open-weight, instruction-tuned model from Google, built as a Mixture-of-Experts so only a fraction of its parameters activate on each token. It handles image, video and text input, and is priced as one of the cheaper multimodal options available through an API.

GoogleVerified 24/09/2026Released 03/04/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 231 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.09

field median $0.43

Output / 1M tokens

$0.30

field median $1.80

Context

262K

field median 500K

02 / overview

What Gemma 4 26B A4B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Gemma 4 26B A4B is an instruction-tuned multimodal model from Google DeepMind, intended for teams that want to run or call a small model without giving up image and video understanding. The sparse Mixture-of-Experts design means a minority of its total parameters are used per token, which is what keeps the serving cost down. It is not a frontier reasoning model and it is not the pick for tasks where you would otherwise reach for Google's largest commercial models.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Gemma 4 26B A4B first appeared in April 2026, as part of Google's Gemma line of open models rather than the Gemini commercial family. Choose it when the work is high volume, latency sensitive, or needs to run somewhere you control, and multimodal input is part of the job. It is the wrong pick for deep, multi step reasoning or anything where you need the strongest available answer per call rather than the cheapest acceptable one.

How you reach it

The API, the apps it powers, and what its limits let you do.

Gemma 4 26B A4B is reached through API providers serving Gemma weights, and as open weights it can also be hosted on your own infrastructure. The context window runs to hundreds of thousands of tokens, and output length is unusually generous for a model this small, so long document sets, screen recordings or video clips can go in and long structured output can come back in one pass. Image, video and text all arrive as input in the same request.

Why it matters

What changes because this exists, or why it does not.

Gemma 4 26B A4B makes multimodal work that was previously priced as a premium feature cheap enough to run across an entire content or media library rather than on a sample. Tagging every product image, summarising every support video, or classifying a long backlog of mixed documents becomes a background job instead of a budget conversation. For teams already running small open models, the addition here is video and a large context window at roughly the same operating cost.

Follows Gemma 4 31B, and is the newest in its line. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Gemma 4 26B A4B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.09per 1M tokens2026-09-24Check
Output$0.30per 1M tokens2026-09-24Check

Gemma 4 26B A4B sits at the low end of the market on both input and output, which is the direct consequence of activating only part of the model per token. Compared with the flagship commercial models it is cheap enough that volume stops being the constraint on what you point it at, and self hosting removes per token cost entirely at the price of running the infrastructure.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$3.30
A busy support assistant200M tokens40M tokens$30.00
A document pipeline1000M tokens100M tokens$120.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window262,144 tokens
Maximum output235,929 tokens
Modalitiesimage, text, video
Released03/04/2026
StatusCurrent
Catalogue identifiergoogle/gemma-4-26b-a4b-it

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

← Gemma 4 31B

Came after

Nothing newer in this line yet.

  1. 01Gemma 3 27B12/03/2025
  2. 02Gemma 3 12B13/03/2025
  3. 03Gemma 3 4B13/03/2025
  4. 04Gemma 4 31B02/04/2026
  5. 05Gemma 4 26B A4B03/04/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap multimodal models like this one mean more of the web gets read, summarised and indexed by machines, including the parts of your site that are images, diagrams and video. If your positioning, pricing and proof only exist inside screenshots or a product demo video, it is now plausible that an engine will extract it, so make sure what it extracts is what you would have said in text. The practical move is to keep the claims in your visual assets consistent with your written pages.

Where buyers meet this model

Buyers rarely meet Gemma 4 26B A4B by name. They meet it inside products built on it, where it is doing the cheap, high volume work behind a chat box, a search feature or a media tagging pipeline, and they meet it directly through hosted API providers or self hosted deployments if they are engineers.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Gemma 4 26B A4B open weight?
Yes. It is part of Google's Gemma line of open models, so it can be self hosted as well as called through hosted API providers.
What can Gemma 4 26B A4B take as input?
Text, images and video. That combination at its price point is the main reason to choose it over a text only small model.
Why is it cheap to run?
It is a Mixture-of-Experts model, so only a small share of its total parameters activate for each token processed. You pay for the active path rather than the full parameter count.
When should I not use Gemma 4 26B A4B?
When the task needs the strongest possible answer rather than the cheapest acceptable one. For hard multi step reasoning, reach for a larger frontier model instead.
Surge45°

Is Gemma 4 26B A4B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.