See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Qwen model°

Qwen3 8B

Qwen3 8B is a dense, text-only open-weight language model from Qwen's Qwen3 series, built to handle both reasoning-heavy work and ordinary dialogue in one small model.

QwenVerified 24/09/2026Released 28/04/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.12

field median $0.30

Output / 1M tokens

$0.46

field median $1.25

Context

131K

field median 524K

02 / overview

What Qwen3 8B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Qwen3 8B is a causal language model of roughly eight billion parameters that can switch between a "thinking" mode for maths and step-by-step problems and a faster mode for normal conversation. The job it was built for is running reasoning workloads at small-model cost, on your own hardware if you want it there. It handles text only, so no image, audio or document-vision work.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Qwen3 8B first appeared in April 2025 as one of the smaller dense members of the Qwen3 family. It is the right pick when you want Qwen3 behaviour at a size you can self-host or run cheaply at volume, and when the task benefits from optional step-by-step reasoning. It is the wrong pick for anything needing images or long generated outputs, and for the hardest reasoning tasks, where the larger Qwen3 models earn their extra cost.

How you reach it

The API, the apps it powers, and what its limits let you do.

Qwen3 8B is reached through the Qwen API and through the hosting providers that serve open Qwen weights, and the weights themselves can be run on your own infrastructure. Its context window is large enough to hold long documents, transcripts or a full retrieval payload in a single call, while the output ceiling keeps it pointed at answers, summaries and structured responses rather than long-form drafting. Mode switching is handled per request, so the same deployment serves both a chat assistant and a reasoning step in a pipeline.

Why it matters

What changes because this exists, or why it does not.

The practical change is that reasoning-style generation stops being something you reserve for a frontier model. A team can put a thinking-mode step inside a high-volume pipeline, classification, extraction, query rewriting, answer checking, and absorb the cost, or run the whole thing in their own environment where data cannot leave. For anyone already using small open models, it is an incremental upgrade rather than a new category.

Follows Qwen3 32B. Superseded by Qwen3.7 Flash. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Qwen3 8B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.12per 1M tokens2026-09-24Check
Output$0.46per 1M tokens2026-09-24Check

Qwen3 8B sits at the cheap end of the hosted market, with output priced at a few multiples of input, as is standard. Running it at volume costs a fraction of a frontier model, and because the weights are open, the hosted price is effectively a ceiling rather than the only option.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$4.62
A busy support assistant200M tokens40M tokens$41.60
A document pipeline1000M tokens100M tokens$162.50

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output8,192 tokens
Modalitiestext
Released28/04/2025
StatusCurrent
Catalogue identifierqwen/qwen3-8b

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Qwen

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

Qwen3 32B
  1. 01Qwen3 14B28/04/2025
  2. 02Qwen3 235B A22B28/04/2025
  3. 03Qwen3 30B A3B28/04/2025
  4. 04Qwen3 32B28/04/2025
  5. 05Qwen3 8B28/04/2025
  6. 06Qwen3.7 Flash27/07/2026
  7. 07Qwen3.8 2.4T A95B12/08/2026
  8. 08Qwen3.8 27B14/08/2026
  9. 09Qwen3.8 Flash26/08/2026
  10. 10Qwen3.8 Max (0902)03/09/2026
  11. 11Qwen3.8 Omni Flash21/09/2026
  12. 12Qwen3.8 Max Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Models in this class are where a lot of production retrieval and summarisation actually runs, so the way your content is written matters as much as which frontier model is in the headlines. A smaller model has less spare capacity to reconcile vague or scattered claims, so it leans harder on clean, self-contained statements of what you do, who you serve and what you cost. If your positioning only makes sense after reading three pages, it will not survive the summarisation step.

Where buyers meet this model

Buyers rarely meet Qwen3 8B by name. They meet it inside products built on it: support assistants, in-app search, summarisation features and internal tools where a vendor chose a cheap self-hostable model, plus the Qwen API and third-party inference platforms that serve open weights.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Can Qwen3 8B read images or documents?
No. Qwen3 8B is text only. Anything visual has to be converted to text before it reaches the model.
What is thinking mode?
It is a switchable behaviour where the model works through a problem step by step before answering, intended for maths and reasoning tasks. The alternative mode is faster and aimed at ordinary dialogue.
When should I use a larger Qwen3 model instead?
When the task is genuinely hard reasoning, or when you need longer generated outputs than this model's response ceiling allows.
Can I run Qwen3 8B myself?
Yes. It is an open-weight dense model at a size most teams can host, which is the main reason to choose it over a closed model of similar cost.
Surge45°

Is Qwen3 8B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.