See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
OpenAI model°

o3

o3 is OpenAI's reasoning model, released in April 2025, built for maths, science, coding and visual reasoning, and available through the OpenAI API and ChatGPT.

OpenAIVerified 24/09/2026Released 16/04/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$2.00

field median $0.30

Output / 1M tokens

$8.00

field median $1.25

Context

200K

field median 524K

02 / overview

What o3 is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

o3 is a general reasoning model rather than a chat model tuned for speed. OpenAI positions it as its strongest all-rounder across domains, with maths, science, coding and visual reasoning as the headline jobs, and technical writing and instruction-following close behind. It is not the right tool for high-volume, low-stakes work where a lighter model would answer just as well for less.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

o3 first appeared in April 2025 as the senior member of OpenAI's o-series reasoning line. Reach for it when a task genuinely needs working-through, a technical spec, a proof, a debugging session across a large codebase, or a diagram that has to be read properly. When the job is summarising, classifying or drafting at volume, a cheaper sibling will do the same work without the reasoning overhead.

How you reach it

The API, the apps it powers, and what its limits let you do.

o3 is reached through the OpenAI API and sits behind the model picker in ChatGPT. It takes text, images and files in the same request, so you can hand it a screenshot, a PDF spec and a question together, and its context window is large enough to hold a long codebase or a full document set in one pass. The generous output ceiling means it can return a complete technical document or a long reasoning trace without being cut short.

Why it matters

What changes because this exists, or why it does not.

o3 made deep reasoning over mixed inputs a normal API call rather than a workaround. Teams that used to split a task into a vision step, a retrieval step and a reasoning step can now send the diagram, the document and the question together and get one worked answer, which removes a lot of orchestration code from technical and scientific pipelines.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

o3 API pricingSurge45°
ChargePriceUnitRead onSource
Input$2.00per 1M tokens2026-09-24Check
Output$8.00per 1M tokens2026-09-24Check

o3 costs meaningfully more per call than the lightweight chat models, and output is charged at several times the input rate, which matters because reasoning models produce long answers. For a frontier reasoning model it is priced to be used routinely rather than saved for special occasions, but it is still worth routing simple requests elsewhere.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$80.00
A busy support assistant200M tokens40M tokens$720.00
A document pipeline1000M tokens100M tokens$2,800.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window200,000 tokens
Maximum output100,000 tokens
Modalitiesimage, text, file
Released16/04/2025
StatusCurrent
Catalogue identifieropenai/o3

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from OpenAI

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Nothing else in this line yet.

A line is worked out from the naming and the release dates across every model page we hold. It fills in as the provider ships successors, or as we pick up the models that came before this one.

How we decide what counts as evidence

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

When a buyer asks a considered, technical question about your category, a reasoning model like o3 will work through the comparison rather than repeat the first source it finds, so thin marketing pages carry less weight than documentation, pricing detail and clear specification. Because it reads images and files alongside text, the material in your PDFs, architecture diagrams and spec sheets is now part of what gets cited, not just your web copy. The practical move is to make the technical substance of your product legible and consistent everywhere it appears.

Where buyers meet this model

Buyers meet o3 inside ChatGPT when they select a reasoning model for a harder question, and engineering teams meet it through the OpenAI API when they build research, coding or document-analysis features. It also sits under third-party products that route technical queries to a stronger model, so a buyer may be reading an o3 answer inside a vendor's own interface without ever seeing the name.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

What is o3 best at?
OpenAI describes o3 as a well-rounded and powerful model across domains, setting a new standard for maths, science, coding and visual reasoning, and also strong at technical writing and instruction-following.
Can o3 read images and documents?
Yes. o3 accepts image, text and file inputs, so you can send a screenshot or diagram alongside a document and a question in the same request.
When should I use something other than o3?
For high-volume, straightforward work such as classification, summarising or routine drafting, a lighter and cheaper model will usually match the result without the cost of a reasoning model.
Surge45°

Is o3 recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.