See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Z.ai model°

GLM 4.5 Air

GLM 4.5 Air is the lightweight member of Z.ai's GLM 4.5 family, a text-only model built for agent-centric applications where the work is tool calls and multi-step tasks rather than open-ended conversation.

Z.aiVerified 24/09/2026Released 25/07/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.13

field median $0.30

Output / 1M tokens

$0.85

field median $1.25

Context

131K

field median 524K

02 / overview

What GLM 4.5 Air is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

GLM 4.5 Air is a compact text model aimed at agent workloads: calling tools, working through steps, producing structured output that another system consumes. It is the smaller sibling in the GLM 4.5 line, trading parameter count for cost and throughput. It does not handle images, audio or any other modality, so anything involving documents as pictures or screen understanding needs a different model.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

GLM 4.5 Air first appeared in July 2025 alongside the wider GLM 4.5 release, positioned below the flagship as the cheaper option for the same agent-shaped jobs. Reach for it when you are running a high volume of similar agent steps and the per-call cost matters more than the last increment of reasoning quality. When a task is genuinely hard, single-shot and expensive to get wrong, the full GLM 4.5 is the better pick.

How you reach it

The API, the apps it powers, and what its limits let you do.

GLM 4.5 Air is reached through the Z.ai API and through the usual third-party model routers that carry open-weight families, so it slots into an existing agent framework without much rework. The context window is large enough to hold a long tool-call history, a set of retrieved documents and a detailed system prompt at the same time, and the generous output ceiling means it can write a long file or a full structured response in one pass rather than being stitched together across calls.

Why it matters

What changes because this exists, or why it does not.

A cheap model that holds a long agent trace changes what you are willing to run repeatedly. Loops that were uneconomic at flagship prices, checking every page of a site, classifying every ticket, retrying a plan three ways, become routine. Nothing here is novel on its own, the significance is that agent-shaped work moves into the price band where you stop counting calls.

Superseded by GLM 4.5V. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

GLM 4.5 Air API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.13per 1M tokens2026-09-24Check
Output$0.85per 1M tokens2026-09-24Check

GLM 4.5 Air sits at the low end of the market, cheap enough that input cost effectively disappears from most budgets and output remains the only line worth watching. Against Western flagship models it is an order of magnitude less expensive, which is the whole argument for using it in high-volume agent loops.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$6.85
A busy support assistant200M tokens40M tokens$60.00
A document pipeline1000M tokens100M tokens$215.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output98,304 tokens
Modalitiestext
Released25/07/2025
StatusCurrent
Catalogue identifierz-ai/glm-4.5-air

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Z.ai

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

The first we track in this line.

Came after

GLM 4.5V
  1. 01GLM 4.5 Air25/07/2025
  2. 02GLM 4.5V11/08/2025
  3. 03GLM 4.630/09/2025
  4. 04GLM 4.6V08/12/2025
  5. 05GLM 5.318/08/2026
  6. 06GLM Latest19/08/2026
  7. 07GLM 5.3 Flash26/08/2026
  8. 08GLM 5.3 FlashX18/09/2026
  9. 09GLM 5.3 Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap agent models mean more automated systems reading your site more often, and reading more of it in one pass given the context they can hold. Your documentation, pricing pages and comparison content are now being ingested by loops that run continuously rather than by a person asking one question. Write pages that survive being read out of order and quoted in fragments, because that is how an agent will use them.

Where buyers meet this model

Most buyers meet GLM 4.5 Air indirectly rather than by name, through the Z.ai API or a model router inside a product someone else built. It is a backend choice, so the encounter is usually a support agent, a research tool or an internal assistant that happens to be running on it.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is GLM 4.5 Air multimodal?
No. GLM 4.5 Air handles text only. If your workflow involves images, screenshots or audio, you need a different model in the pipeline.
When should I use GLM 4.5 Air instead of GLM 4.5?
Use Air when you are running the same kind of agent step many times and cost per call is the constraint. Use the full GLM 4.5 when the task is hard, one-off, and the cost of a wrong answer outweighs the saving.
What is GLM 4.5 Air built for?
Agent-centric applications: tool calling, multi-step task execution and structured output that feeds another system. It is a working model inside a pipeline rather than a chat product.
Surge45°

Is GLM 4.5 Air recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.