See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Qwen model°

Qwen3 30B A3B

Qwen3 30B A3B is a text model from Qwen, part of the Qwen3 generation, built as a mixture-of-experts model that keeps only a small share of its parameters active per token. It is aimed at reasoning, multilingual work and agent-style tasks at a cost well below the frontier tier.

QwenVerified 24/09/2026Released 28/04/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.12

field median $0.30

Output / 1M tokens

$0.50

field median $1.25

Context

131K

field median 524K

02 / overview

What Qwen3 30B A3B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Qwen3 30B A3B is a text-only large language model in the Qwen3 series, which spans both dense and mixture-of-experts designs. It was built for reasoning, multilingual handling and agent tasks, meaning tool calls and multi-step chains rather than one-shot completions. It does not read images, audio or video, so anything involving documents as pictures or screenshots needs a different model.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Qwen3 30B A3B first appeared in April 2025 as one of the mixture-of-experts entries in the Qwen3 line-up, sitting between the small dense models and the larger flagship. It is the right pick when you want Qwen3 reasoning behaviour on high-volume work and cannot justify a frontier model per call. It is the wrong pick when the task needs vision, or when you need long generated outputs in a single response, since its output ceiling is modest next to its input capacity.

How you reach it

The API, the apps it powers, and what its limits let you do.

You reach it through an API as a text-in, text-out model, either from Qwen directly or through the aggregators and open-weight hosts that carry the Qwen3 family. The context window is large enough to hold a full documentation set, a long support thread or a codebase slice in one call, while the output limit keeps replies to the length of a long answer or a structured payload rather than a full report. Its agent orientation means it is usually wired into a tool-calling loop rather than used as a bare chat endpoint.

Why it matters

What changes because this exists, or why it does not.

Because only a fraction of the model is active per token, teams get reasoning and tool-use behaviour at a price point that previously bought a much weaker model, which changes the maths on running classification, extraction or retrieval grading across every record rather than a sample. The multilingual coverage means one model can serve several markets instead of separate per-language routing. For English-only, low-volume chat, it is unremarkable, plenty of models do that.

Follows Qwen3 235B A22B. Superseded by Qwen3 32B. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Qwen3 30B A3B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.12per 1M tokens2026-09-24Check
Output$0.50per 1M tokens2026-09-24Check

Running Qwen3 30B A3B costs a small fraction of what frontier models charge, which is the point of the sparse design, you pay for the active slice rather than the full parameter count. At this level, cost stops being the constraint on batch and always-on workloads, and engineering time becomes the real budget line.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$4.90
A busy support assistant200M tokens40M tokens$44.00
A document pipeline1000M tokens100M tokens$170.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output16,384 tokens
Modalitiestext
Released28/04/2025
StatusCurrent
Catalogue identifierqwen/qwen3-30b-a3b

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Qwen

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Qwen3 14B28/04/2025
  2. 02Qwen3 235B A22B28/04/2025
  3. 03Qwen3 30B A3B28/04/2025
  4. 04Qwen3 32B28/04/2025
  5. 05Qwen3 8B28/04/2025
  6. 06Qwen3.7 Flash27/07/2026
  7. 07Qwen3.8 2.4T A95B12/08/2026
  8. 08Qwen3.8 27B14/08/2026
  9. 09Qwen3.8 Flash26/08/2026
  10. 10Qwen3.8 Max (0902)03/09/2026
  11. 11Qwen3.8 Omni Flash21/09/2026
  12. 12Qwen3.8 Max Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap reasoning models change the economics of AI answer engines, because a system that can afford to run a model over every retrieved page will read more sources before it answers, not fewer. That rewards content that holds up when it is actually read rather than skimmed: clear claims, named products, specifics a model can lift and attribute. It also means your visibility is no longer decided only in English, so a brand with thin or machine-translated pages in other languages will be quietly absent from those answers.

Where buyers meet this model

Buyers meet Qwen3 30B A3B mostly through the API rather than a consumer app, either via Qwen's own platform or through the inference providers and open-weight hosts that serve the Qwen3 family. It also shows up inside other people's products, as the layer doing retrieval grading, summarising or agent steps in tools that never name the model on screen. In Chinese-language and wider Asian markets, Qwen models sit behind a meaningful share of assistant and search experiences.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Can Qwen3 30B A3B read images or PDFs?
No. It is a text-only model. Anything arriving as a scan, screenshot or image-based PDF needs to be converted to text first, or handled by a model with vision input.
Is it suitable for agent workflows?
Yes, the Qwen3 generation is built with agent tasks in mind, so it is typically used inside a tool-calling loop rather than as a plain chat endpoint. The large context window helps when an agent accumulates tool results across many steps.
How does it compare with the larger Qwen3 models?
It sits below the flagship in the family and is chosen when you want Qwen3 behaviour at high volume rather than maximum capability per call. For the hardest single reasoning tasks, a larger sibling is the better fit.
Why is the output limit lower than the context window?
It can take in far more than it can write out in one response, which suits summarising, extraction and grading over long inputs. If you need a long generated document, plan to produce it in sections across multiple calls.
Surge45°

Is Qwen3 30B A3B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.