See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Z.ai model°

GLM 4.6

GLM 4.6 is a text-only large language model from Z.ai, the follow-on to GLM-4.5, built for long-context work at a low per-token cost.

Z.aiVerified 24/09/2026Released 30/09/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.43

field median $0.30

Output / 1M tokens

$1.75

field median $1.25

Context

205K

field median 524K

02 / overview

What GLM 4.6 is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

GLM 4.6 is a general-purpose text model, not a multimodal one: it reads and writes text, with no image, audio or video input. Z.ai positions it as an incremental upgrade over GLM-4.5, with the headline change being a larger context window for handling more complex inputs in a single pass. If your job needs vision or speech, this is the wrong family.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

GLM 4.6 first appeared at the end of September 2025 as the current generation of Z.ai's GLM line, replacing GLM-4.5 as the default choice within that family. It is the right pick when you are already working with GLM models and want the longer context without changing providers, or when cost per token is the binding constraint. It is the wrong pick when you need multimodal input or very long generated outputs, since the output ceiling is modest.

How you reach it

The API, the apps it powers, and what its limits let you do.

GLM 4.6 is reached through Z.ai's API and through the aggregators that resell it, and it is text in, text out. The expanded context window means you can put a large document set, a long transcript or a sizeable codebase into one request, but the output limit keeps each response comparatively short, so long-form generation has to be chunked across calls.

Why it matters

What changes because this exists, or why it does not.

The practical change is that long-context work gets cheap. Feeding an entire contract, a quarter of support tickets or a full documentation set into a single prompt was previously something you rationed on cost, and GLM 4.6 makes it a routine call, which suits batch classification, extraction and retrieval-heavy pipelines more than it suits flagship chat.

Follows GLM 4.5V. Superseded by GLM 4.6V. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

GLM 4.6 API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.43per 1M tokens2026-09-24Check
Output$1.75per 1M tokens2026-09-24Check

GLM 4.6 sits at the budget end of the market, cheap enough that large-context prompts stop being something you budget around. Against the frontier models from the larger US labs it costs a small fraction per call, which is the main reason teams pick it for high-volume, repetitive text work.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$17.35
A busy support assistant200M tokens40M tokens$156.00
A document pipeline1000M tokens100M tokens$605.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window204,800 tokens
Maximum output16,384 tokens
Modalitiestext
Released30/09/2025
StatusCurrent
Catalogue identifierz-ai/glm-4.6

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Z.ai

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

GLM 4.5V

Came after

GLM 4.6V
  1. 01GLM 4.5 Air25/07/2025
  2. 02GLM 4.5V11/08/2025
  3. 03GLM 4.630/09/2025
  4. 04GLM 4.6V08/12/2025
  5. 05GLM 5.318/08/2026
  6. 06GLM Latest19/08/2026
  7. 07GLM 5.3 Flash26/08/2026
  8. 08GLM 5.3 FlashX18/09/2026
  9. 09GLM 5.3 Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Models like GLM 4.6 mean your content is being read and summarised at a price point where nobody thinks twice about ingesting your entire site in one go. That rewards source material that holds up when read whole and in context, consistent product naming, clear pricing pages, documentation that answers questions directly, rather than pages tuned for a snippet. If a builder wires a cheap long-context model into their answer layer, the brands that get cited are the ones whose full text is unambiguous.

Where buyers meet this model

Most buyers meet GLM 4.6 through the API rather than a consumer app, either directly from Z.ai or via a model router where it sits alongside Western models as a lower-cost option. It also turns up inside products whose builders chose it on price, so a buyer may be reading its output without ever seeing the name.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Can GLM 4.6 handle images?
No. GLM 4.6 is text only, so images, audio and video are out of scope. You would need a separate multimodal model for anything involving visual or spoken input.
How is GLM 4.6 different from GLM-4.5?
Z.ai describes GLM 4.6 as a generational improvement over GLM-4.5, with the main stated change being a substantially larger context window so the model can work with more complex and longer inputs in a single request.
What is GLM 4.6 best used for?
High-volume text work where the input is long and the output is short: extraction, classification, summarisation and retrieval-heavy pipelines. Its output ceiling makes it less suited to generating long documents in one call.
Who makes GLM 4.6?
GLM 4.6 comes from Z.ai and is the current generation of its GLM model line, first seen at the end of September 2025.
Surge45°

Is GLM 4.6 recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.