See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
OpenAI model°

GPT-5.1-Codex

GPT-5.1-Codex is OpenAI's coding-specialised version of GPT-5.1, built for software engineering work rather than general chat. It handles both interactive development sessions and long stretches of independent work on complex engineering tasks.

OpenAIVerified 24/09/2026Released 13/11/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$1.25

field median $0.30

Output / 1M tokens

$10.00

field median $1.25

Context

400K

field median 524K

02 / overview

What GPT-5.1-Codex is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

GPT-5.1-Codex is a variant of GPT-5.1 tuned for software engineering and coding workflows. The job it is built for is writing, reading and changing code, either alongside a developer in a session or running on its own through a long task. It is not the model to reach for when you want general writing, research or open-ended conversation, that is what the base GPT-5.1 is for.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

GPT-5.1-Codex first appeared in November 2025, as the coding-specialised sibling in the GPT-5.1 family. It is the right pick when the work is a codebase, particularly when a task needs to run for a long time without a human turning every crank. It is the wrong pick for anything outside engineering, where the general GPT-5.1 will behave better and you gain nothing from the specialisation.

How you reach it

The API, the apps it powers, and what its limits let you do.

GPT-5.1-Codex is reached through the OpenAI API and is the model behind OpenAI's Codex coding tooling. It accepts text and images, so screenshots, diagrams and design mockups can go into a prompt alongside code. The context window is large enough to hold a substantial slice of a repository, plus logs and diffs, in a single pass, and the output ceiling is high enough to return large multi-file changes rather than fragments.

Why it matters

What changes because this exists, or why it does not.

The practical change is duration. Agentic coding work that previously had to be chopped into small supervised steps can be handed over as one long task, with enough context to keep the whole picture in view and enough output room to return the full change set. For teams already running coding agents, this is a step up in how much you can delegate in one go rather than a new category of thing.

Follows GPT-5.1. Superseded by GPT-5.1-Codex-Max. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

GPT-5.1-Codex API pricingSurge45°
ChargePriceUnitRead onSource
Input$1.25per 1M tokens2026-09-24Check
Output$10.00per 1M tokens2026-09-24Check

It sits at the same rate as the general GPT-5.1 line, so choosing the coding variant costs no premium over the standard model. Against the wider field it is mid-market, well above the small fast models used for classification and routing, and well below the top-tier reasoning models, though long agentic runs on a large context will consume tokens quickly whatever the unit rate.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$75.00
A busy support assistant200M tokens40M tokens$650.00
A document pipeline1000M tokens100M tokens$2,250.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window400,000 tokens
Maximum output128,000 tokens
Modalitiestext, image
Released13/11/2025
StatusCurrent
Catalogue identifieropenai/gpt-5.1-codex

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from OpenAI

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01GPT-4.114/04/2025
  2. 02GPT-4.1 Mini14/04/2025
  3. 03GPT-4.1 Nano14/04/2025
  4. 04GPT-507/08/2025
  5. 05GPT-5 Mini07/08/2025
  6. 06GPT-5 Nano07/08/2025
  7. 07GPT-5 Pro06/10/2025
  8. 08GPT-5.113/11/2025
  9. 09GPT-5.1-Codex13/11/2025
  10. 10GPT-5.1-Codex-Max04/12/2025
  11. 11GPT-5.210/12/2025
  12. 12GPT-5.2 Pro10/12/2025
  13. 13GPT-6 Astra04/09/2026
  14. 14GPT-6 Astra Pro04/09/2026
  15. 15GPT-6 Luna22/09/2026
  16. 16GPT-6 Luna Pro22/09/2026
  17. 17GPT-6 Sol22/09/2026
  18. 18GPT-6 Sol Pro22/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

A coding-specialised model does not answer buyer questions, so it will not cite your marketing pages. Where it matters for a SaaS brand is developer-facing content: if you sell an API, an SDK or anything a developer integrates, your docs, code samples and error messages are now being read by an agent working through a long task, not skimmed by a human. Clear, complete, copy-pasteable documentation is what gets your product used correctly inside those sessions.

Where buyers meet this model

Buyers meet GPT-5.1-Codex mostly through developer tooling rather than a consumer chat app: OpenAI's Codex products and the API, plus any IDE extension, CI job or coding agent a team has wired up to it. It is not a model your prospects will encounter while asking a question in a chatbot, so it sits outside the usual AI search surfaces.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

How is GPT-5.1-Codex different from GPT-5.1?
It is the same family, tuned for software engineering and coding workflows. Use it when the task is a codebase and use the general GPT-5.1 for everything else.
Can GPT-5.1-Codex work from screenshots?
Yes. It takes text and images, so screenshots, mockups and diagrams can be passed in alongside code.
Is GPT-5.1-Codex suitable for long-running agent tasks?
That is one of the two things it is designed for, along with interactive development sessions. It is built to run independently through complex engineering tasks rather than needing a human at every step.
Does GPT-5.1-Codex affect whether our SaaS gets cited in AI answers?
Not directly, it is a developer tool rather than an answer engine. It matters indirectly if you sell to developers, because your documentation and code samples are what the model reads when someone builds against your product.
Surge45°

Is GPT-5.1-Codex recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.