See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Z.ai model°

GLM 4.6V

GLM 4.6V is a multimodal model from Z.ai built for high-fidelity visual understanding, reading images, video and document pages alongside text in a single long-context request.

Z.aiVerified 24/09/2026Released 08/12/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.30

field median $0.30

Output / 1M tokens

$0.90

field median $1.25

Context

131K

field median 524K

02 / overview

What GLM 4.6V is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

GLM 4.6V is a large multimodal model whose job is visual understanding and long-context reasoning over mixed media: images, documents and video together with text. It is built to handle complex page layouts rather than to act as a plain chat assistant, so the natural work is extraction, description and reasoning about what a page or a frame actually contains. It is not a text-only writing model, and there is nothing in Z.ai's description to suggest it generates images or video of its own.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

GLM 4.6V was first seen in December 2025 and sits in Z.ai's GLM family as the vision-capable member, the one you reach for when the input is not purely text. Choose it when a task depends on seeing a document, a screenshot or a video frame, and the material runs long enough that chunking would break the reasoning. If the work is text in and text out, a text-only sibling is the more sensible pick.

How you reach it

The API, the apps it powers, and what its limits let you do.

GLM 4.6V is reached through Z.ai's API, which takes image, video and text input in the same call. Its context window is large enough to hold a long document set or a run of frames at once, so you can ask questions across a whole report rather than page by page, and the output ceiling is generous enough for structured extraction results and long summaries rather than short answers only. That combination is what makes document pipelines, screenshot review and media tagging practical on a single model.

Why it matters

What changes because this exists, or why it does not.

Work that used to need an OCR step, a layout parser and a separate reasoning model can now sit behind one call, which removes the seams where meaning usually gets lost. Video as a first-class input matters too: teams handling recordings, product demos or screen captures no longer have to reduce them to stills and captions before a model can reason about them. The combination of visual fidelity and a long window is the part that was awkward before, since most vision models made you trade one for the other.

Follows GLM 4.6. Superseded by GLM 5.3. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

GLM 4.6V API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.30per 1M tokens2026-09-24Check
Output$0.90per 1M tokens2026-09-24Check

GLM 4.6V is priced at the budget end of the multimodal field, closer to a small text model than to the frontier vision systems it competes with on task. That makes high-volume document and video processing affordable to run continuously rather than as a one-off batch, which is usually the constraint that kills these pipelines.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$10.50
A busy support assistant200M tokens40M tokens$96.00
A document pipeline1000M tokens100M tokens$390.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output32,768 tokens
Modalitiesimage, text, video
Released08/12/2025
StatusCurrent
Catalogue identifierz-ai/glm-4.6v

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Z.ai

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

GLM 4.6

Came after

GLM 5.3
  1. 01GLM 4.5 Air25/07/2025
  2. 02GLM 4.5V11/08/2025
  3. 03GLM 4.630/09/2025
  4. 04GLM 4.6V08/12/2025
  5. 05GLM 5.318/08/2026
  6. 06GLM Latest19/08/2026
  7. 07GLM 5.3 Flash26/08/2026
  8. 08GLM 5.3 FlashX18/09/2026
  9. 09GLM 5.3 Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

A multimodal model at this price makes it cheap for someone to read your product visually, so screenshots, pricing pages, dashboards and demo videos become citable source material, not just your written copy. If your differentiators only exist inside an interface or a video, they are now readable, and if your pages are image-heavy with the substance locked in graphics, that substance is finally legible to an answer engine. The practical move is to make sure what a model sees in your screenshots and diagrams agrees with what your text claims.

Where buyers meet this model

Buyers meet GLM 4.6V mostly through the API, embedded inside someone else's product rather than as a named chat companion. It shows up in document review tools, support and onboarding flows that read user screenshots, and internal search over scanned or image-heavy archives. Z.ai's own GLM surfaces are the other place it appears directly.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

What is GLM 4.6V used for?
GLM 4.6V is used for visual understanding and long-context reasoning across images, documents, video and text. Typical work includes reading complex page layouts, extracting structured data from scanned or image-based documents, and answering questions about screenshots and recorded media.
Does GLM 4.6V handle video?
Yes. GLM 4.6V accepts video as an input modality alongside images and text, so recordings and screen captures can be reasoned about directly rather than reduced to stills first.
Is GLM 4.6V good for long documents?
Yes. Its context window is large enough to hold a long document or document set in one request, which means questions can be answered across the whole thing instead of chunk by chunk. Z.ai specifically describes it as handling complex page layouts.
How does GLM 4.6V compare with other multimodal models on cost?
It sits at the inexpensive end of the multimodal market, which makes continuous, high-volume image and video processing viable rather than something you run in occasional batches.
Surge45°

Is GLM 4.6V recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.