See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Qwen model°

Qwen3 32B

Qwen3 32B is an open-weight, text-only language model from Qwen, a dense model in the Qwen3 series that can switch between a "thinking" mode for step-by-step reasoning and a faster mode for ordinary dialogue.

QwenVerified 24/09/2026Released 28/04/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.08

field median $0.30

Output / 1M tokens

$0.28

field median $1.25

Context

131K

field median 524K

02 / overview

What Qwen3 32B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Qwen3 32B is a dense causal language model built to handle both complex reasoning and everyday conversational work in one set of weights, with a mode switch rather than two separate models. It takes text in and produces text out, so it is not the model for reading screenshots, PDFs as images, audio or video. It sits in the mid-size band of the Qwen3 family, meant for teams that want reasoning behaviour without running a frontier-scale model.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Qwen3 32B first appeared at the end of April 2025 as part of the Qwen3 release. It is the right pick when you want reasoning quality from a self-hostable dense model and are willing to trade some headroom for cost and control, and it is the wrong pick when the job needs images or audio, or when you need very long generated outputs in a single pass.

How you reach it

The API, the apps it powers, and what its limits let you do.

Qwen3 32B is reached through the Qwen API and through the usual third-party inference hosts that carry open Qwen weights, and the weights themselves can be run on your own hardware. Its context window is large enough to hold a full documentation set, a long support thread or a codebase slice in one prompt, though the output ceiling keeps it pointed at answers and summaries rather than book-length generation. The thinking mode is toggled per request, so the same deployment can serve slow analytical work and quick chat turns.

Why it matters

What changes because this exists, or why it does not.

The practical change is that reasoning-style output stops being something you only buy from a hosted frontier model. A team can run Qwen3 32B on its own infrastructure, turn thinking on for the small share of requests that need it, and leave it off for the rest, which was awkward when reasoning and chat meant maintaining two different models with two different prompt formats.

Follows Qwen3 30B A3B. Superseded by Qwen3 8B. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Qwen3 32B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.08per 1M tokens2026-09-24Check
Output$0.28per 1M tokens2026-09-24Check

Qwen3 32B is among the cheapest models a serious application would consider, with input costing a small fraction of what hosted frontier models charge and output not much more. At that level, the cost question shifts from per-call economics to whether self-hosting is worth the operational effort at your volume.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$3.00
A busy support assistant200M tokens40M tokens$27.20
A document pipeline1000M tokens100M tokens$108.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output16,384 tokens
Modalitiestext
Released28/04/2025
StatusCurrent
Catalogue identifierqwen/qwen3-32b

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Qwen

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came after

Qwen3 8B
  1. 01Qwen3 14B28/04/2025
  2. 02Qwen3 235B A22B28/04/2025
  3. 03Qwen3 30B A3B28/04/2025
  4. 04Qwen3 32B28/04/2025
  5. 05Qwen3 8B28/04/2025
  6. 06Qwen3.7 Flash27/07/2026
  7. 07Qwen3.8 2.4T A95B12/08/2026
  8. 08Qwen3.8 27B14/08/2026
  9. 09Qwen3.8 Flash26/08/2026
  10. 10Qwen3.8 Max (0902)03/09/2026
  11. 11Qwen3.8 Omni Flash21/09/2026
  12. 12Qwen3.8 Max Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap open-weight reasoning models mean more products can afford to read your documentation at length before answering, so long, well-structured source pages get used rather than skimmed. Because Qwen3 32B is text-only, anything you explain in a diagram, screenshot or video is invisible to it, and the written version is the only version it will cite. Being retrievable in clean text, with claims stated plainly near the top of a page, matters more here than in the multimodal tier.

Where buyers meet this model

Buyers meet Qwen3 32B mostly through products built on it rather than through a branded consumer app: in-product assistants, support bots and internal research tools where a vendor picked open weights for cost or data-residency reasons. It is also widely available across inference marketplaces, which is how most developers first try it. If you are a SaaS brand, the model is more likely to be reading your docs inside somebody else's product than answering a consumer directly.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Can Qwen3 32B read images or PDFs?
No. Qwen3 32B is text-only, so images, scanned documents, audio and video need to be converted to text before they reach it.
What is thinking mode in Qwen3 32B?
It is a switchable behaviour in the model that runs step-by-step reasoning before answering. You can turn it on for hard analytical requests and leave it off for ordinary dialogue, using the same model in both cases.
Is Qwen3 32B suitable for long documents?
Yes for input. Its context window comfortably holds long documents or a set of them, but the output limit means it is better suited to producing answers, extracts and summaries than very long generated text.
Surge45°

Is Qwen3 32B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.