See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
Anthropic model°

Claude Sonnet 4

Claude Sonnet 4 is Anthropic's mid-tier Claude model, first seen in May 2025, built as the everyday workhorse for coding and reasoning tasks with a large context window and image, text and file input.

AnthropicVerified 24/09/2026Released 22/05/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$3.00

field median $0.30

Output / 1M tokens

$15.00

field median $1.25

Context

200K

field median 524K

02 / overview

What Claude Sonnet 4 is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Claude Sonnet 4 is a general-purpose reasoning and coding model, positioned as the step up from Sonnet 3.7 rather than a research-grade frontier release. Anthropic points to its coding performance, with a state-of-the-art SWE-bench result of 72.7%, and to tighter precision and controllability, meaning it does what the instruction says instead of improvising around it. It is not a specialist retrieval or embedding model, and it is not the cheapest option for high-volume classification work.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Claude Sonnet 4 arrived in May 2025 as the direct successor to Claude Sonnet 3.7, sitting in the middle of the Claude family between the small and the largest models. It is the right pick when a task needs real reasoning or code changes across a sizeable codebase but does not justify the cost of the top tier. It is the wrong pick for trivial, high-throughput calls where a smaller model would do the same job for less.

How you reach it

The API, the apps it powers, and what its limits let you do.

Claude Sonnet 4 is reached through the Anthropic API and powers the Claude consumer apps. It takes text, images and files in the same request, so you can hand it a screenshot, a spec document and a prompt together, and its context window is large enough to hold a long codebase or a full set of source documents in one pass. The generous maximum output means it can return complete files or long structured documents rather than fragments you have to stitch back together.

Why it matters

What changes because this exists, or why it does not.

The controllability improvement matters more than the benchmark number for anyone building products on it: instructions are followed more literally, so agentic and tool-using workflows need fewer guardrails to stay on track. Teams that previously escalated coding and multi-step reasoning work to a top-tier model can run more of it here, which changes the economics of agents that make many calls per task.

Superseded by Claude Sonnet 4.5. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Claude Sonnet 4 API pricingSurge45°
ChargePriceUnitRead onSource
Input$3.00per 1M tokens2026-09-24Check
Output$15.00per 1M tokens2026-09-24Check

Claude Sonnet 4 sits in the mid-range of the market, meaningfully cheaper to run than frontier-tier models while costing more than the small, fast models aimed at bulk work. Output costs several times more than input, so the practical lever on your bill is keeping responses tight rather than trimming the context you send.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$135.00
A busy support assistant200M tokens40M tokens$1,200.00
A document pipeline1000M tokens100M tokens$4,500.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window200,000 tokens
Maximum output64,000 tokens
Modalitiesimage, text, file
Released22/05/2025
StatusCurrent
Catalogue identifieranthropic/claude-sonnet-4

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from Anthropic

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

The first we track in this line.

  1. 01Claude Sonnet 422/05/2025
  2. 02Claude Sonnet 4.529/09/2025

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Because Sonnet 4 handles a large share of default Claude traffic, it is one of the main models deciding whether your brand appears in an AI answer at all. Its long context means it can read a full documentation page, pricing page or comparison article rather than a snippet, so depth and clarity on your own site are rewarded. Its file and image input also means prospects can drop in a competitor's PDF or a screenshot of your pricing and ask it to compare, which is an evaluation you should assume is happening.

Where buyers meet this model

Buyers meet Claude Sonnet 4 most often inside the Claude apps, where it is the default tier for a large share of everyday conversations, including product research and shortlisting. Developers meet it through the Anthropic API and through the coding tools and IDE assistants that route to it. Either way, when somebody asks Claude which tool to use, this is frequently the model composing the answer.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

How is Claude Sonnet 4 different from Claude Sonnet 3.7?
Anthropic describes it as a significant enhancement over Sonnet 3.7, specifically in coding and reasoning, with improved precision and controllability. In practice that means it follows instructions more closely and holds up better across multi-step tasks.
Is Claude Sonnet 4 good for coding?
Coding is its headline strength. Anthropic reports a state-of-the-art SWE-bench score of 72.7%, and its large context window lets it work across a substantial codebase in a single pass.
Can Claude Sonnet 4 read images and documents?
Yes. It accepts image, text and file input, so you can combine a screenshot, an uploaded document and a written prompt in the same request.
When should you use a larger Claude model instead?
Sonnet 4 is the mid-tier pick, so reach for the top tier when a task needs the deepest available reasoning and the extra cost is justified. For simple, high-volume calls, a smaller and cheaper model is usually the better fit.
Surge45°

Is Claude Sonnet 4 recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.