See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Qwen model°

Qwen3 14B

Qwen3 14B is a dense open-weight language model from Qwen's Qwen3 series, built to handle both step-by-step reasoning and ordinary dialogue in one model by switching between a "thinking" mode and a faster direct mode.

QwenVerified 24/09/2026Released 28/04/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.12

field median $0.30

Output / 1M tokens

$0.24

field median $1.25

Context

131K

field median 524K

02 / overview

What Qwen3 14B is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Qwen3 14B is a text-only causal language model in the mid-size band of the Qwen3 family, sized to run complex reasoning work without the cost of a frontier model. Its defining feature is the ability to switch between a deliberate thinking mode and a straight conversational mode, so one deployment covers both analysis and chat. It is not a multimodal model, so images, audio and video are out of scope.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Qwen3 14B appeared in April 2025 as part of the Qwen3 release. It sits in the middle of the series, above the small models that are cheap but limited on reasoning, and below the larger Qwen3 variants. Pick it when you want reasoning quality on a modest budget, or when you are self-hosting and want something that fits a single sensible GPU. It is the wrong pick when you need vision or audio input, or when the task genuinely needs the largest model in the family.

How you reach it

The API, the apps it powers, and what its limits let you do.

Qwen3 14B is reached through API providers that host the Qwen3 series, and the weights are open so it can also be run on your own infrastructure. It is text in, text out, with a long context window that comfortably holds large documents, long chat histories or sizeable codebases in a single request. The per-request output cap is more modest than the input allowance, so it suits analysing a lot and replying concisely rather than generating book-length output.

Why it matters

What changes because this exists, or why it does not.

The practical change is that reasoning and fast dialogue no longer need two separate models behind one product. A team can route a support conversation and a multi-step analysis task to the same endpoint and toggle the mode, which removes a routing layer and a second set of prompts to maintain. For anyone running models themselves, a dense model at this size is a straightforward deployment rather than a cluster project.

Superseded by Qwen3 235B A22B. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Qwen3 14B API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.12per 1M tokens2026-09-24Check
Output$0.24per 1M tokens2026-09-24Check

Qwen3 14B is priced in the cheap tier of hosted models, with output costing double the input, which is the usual shape. Against frontier chat models it is a fraction of the cost, and because the weights are open the hosted price is really a ceiling, self-hosting at volume takes it lower still.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$3.60
A busy support assistant200M tokens40M tokens$33.60
A document pipeline1000M tokens100M tokens$144.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window131,072 tokens
Maximum output16,384 tokens
Modalitiestext
Released28/04/2025
StatusCurrent
Catalogue identifierqwen/qwen3-14b

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Qwen

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

The first we track in this line.

  1. 01Qwen3 14B28/04/2025
  2. 02Qwen3 235B A22B28/04/2025
  3. 03Qwen3 30B A3B28/04/2025
  4. 04Qwen3 32B28/04/2025
  5. 05Qwen3 8B28/04/2025
  6. 06Qwen3.7 Flash27/07/2026
  7. 07Qwen3.8 2.4T A95B12/08/2026
  8. 08Qwen3.8 27B14/08/2026
  9. 09Qwen3.8 Flash26/08/2026
  10. 10Qwen3.8 Max (0902)03/09/2026
  11. 11Qwen3.8 Omni Flash21/09/2026
  12. 12Qwen3.8 Max Prime23/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Open models like Qwen3 14B get embedded quietly into other companies' assistants and internal tools, so your brand can be cited in answers generated by a model you never see in a logo wall. What it rewards is the same thing every retrieval-backed system rewards: clear, self-contained pages that state what your product does and who it is for, because a mid-size model has less general knowledge to fall back on and leans harder on what it is given. Assume the retrieved page is doing most of the work, and write it accordingly.

Where buyers meet this model

Buyers rarely meet Qwen3 14B by name. They meet it through the API at aggregators and inference hosts that carry the Qwen3 series, inside products whose vendors chose an open model for cost or data-residency reasons, and on self-hosted deployments behind an internal assistant or a support widget.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Qwen3 14B open weight?
Yes. It is part of the openly released Qwen3 series, so it can be run on your own infrastructure as well as called through hosted API providers.
What is thinking mode in Qwen3 14B?
Qwen3 14B can switch between a thinking mode, where it works through a problem step by step before answering, and a direct mode for efficient dialogue. The same model covers both, so you do not need separate reasoning and chat deployments.
Can Qwen3 14B read images or PDFs with scans?
No. It is a text-only model, so anything visual needs to be converted to text before it reaches the model.
When should I use a larger Qwen3 model instead?
When the task needs the strongest reasoning the family offers and the cost difference is not the constraint. Qwen3 14B is the pick when you want solid reasoning at low cost or a model small enough to host comfortably yourself.
Surge45°

Is Qwen3 14B recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.