See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
OpenAI model°

GPT-5.6 Luna

GPT-5.6 Luna is the fast, low-cost model in OpenAI's GPT-5.6 series, built for high-volume and latency-sensitive work such as chat, classification and lightweight agentic workflows.

OpenAIVerified 24/09/2026Released 09/07/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 151 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.20

field median $0.43

Output / 1M tokens

$1.20

field median $1.75

Context

1,050K

field median 524K

02 / overview

What GPT-5.6 Luna is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

GPT-5.6 Luna is a general-purpose text and multimodal model positioned as the cheap, fast option within its series. It was built for work that runs at volume, classification, routing, chat responses and short agent steps, where latency and unit cost matter more than depth. It is not the model to reach for when a task needs sustained, careful reasoning, that is what its larger siblings in the series are for.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

GPT-5.6 Luna was first seen in July 2026 as the fast tier of OpenAI's GPT-5.6 series. Pick it when a task runs thousands of times a day and each run is individually simple: tagging tickets, extracting fields, answering routine questions, driving a small agent loop. Pick a heavier model in the series when correctness on a single hard question matters more than throughput.

How you reach it

The API, the apps it powers, and what its limits let you do.

GPT-5.6 Luna is reached through the OpenAI API and sits behind OpenAI's own consumer and developer surfaces as a fast tier. It takes text, images and files as input, so a request can carry a scanned document or a screenshot alongside the prompt, and its context window is large enough to hold a full document set, a long support history or a sizeable codebase in a single call rather than chunking it. The output ceiling is generous enough for long structured returns, full reports or bulk JSON, not just short replies.

Why it matters

What changes because this exists, or why it does not.

The combination of a very large context window and a low per-token price changes what is worth reading. Work that previously meant retrieval, chunking and a fair amount of plumbing, feeding a model a long document set or a full conversation history, can now be handled by passing the whole thing in and paying very little for it. For teams running classification or agent loops at scale, the cost of being thorough has dropped to the point where it is no longer a design constraint.

Follows GPT-5.2 Pro. Superseded by GPT-5.6 Luna Pro. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

GPT-5.6 Luna API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.20per 1M tokens2026-09-24Check
Output$1.20per 1M tokens2026-09-24Check

GPT-5.6 Luna is priced at the cheap end of the current field, with output costing several times more than input, which is the usual shape for a fast tier. The practical effect is that reading is close to free and writing is where your bill accumulates, so it suits jobs that ingest a lot and return a little, and it is affordable enough to sit on every request rather than being reserved for the hard ones.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$10.00
A busy support assistant200M tokens40M tokens$88.00
A document pipeline1000M tokens100M tokens$320.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window1,050,000 tokens
Maximum output128,000 tokens
Modalitiesfile, image, text
Released09/07/2026
StatusCurrent
Catalogue identifieropenai/gpt-5.6-luna

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from OpenAI

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01GPT-4.114/04/2025
  2. 02GPT-4.1 Mini14/04/2025
  3. 03GPT-4.1 Nano14/04/2025
  4. 04GPT-507/08/2025
  5. 05GPT-5 Mini07/08/2025
  6. 06GPT-5 Nano07/08/2025
  7. 07GPT-5 Pro06/10/2025
  8. 08GPT-5.113/11/2025
  9. 09GPT-5.1-Codex13/11/2025
  10. 10GPT-5.1-Codex-Max04/12/2025
  11. 11GPT-5.210/12/2025
  12. 12GPT-5.2 Pro10/12/2025
  13. 13GPT-5.6 Luna09/07/2026
  14. 14GPT-5.6 Luna Pro09/07/2026
  15. 15GPT-5.6 Sol09/07/2026
  16. 16GPT-5.6 Sol Pro09/07/2026
  17. 17GPT-5.6 Terra09/07/2026
  18. 18GPT-5.6 Terra Pro09/07/2026
  19. 19GPT-6 Astra04/09/2026
  20. 20GPT-6 Astra Pro04/09/2026
  21. 21GPT-6 Luna22/09/2026
  22. 22GPT-6 Luna Pro22/09/2026
  23. 23GPT-6 Sol22/09/2026
  24. 24GPT-6 Sol Pro22/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

A cheap, long-context model with file and image input makes it economically sensible for an AI surface to read your whole documentation set, pricing page and comparison content on a single query rather than skimming a summary. That rewards brands whose material is complete, structured and machine-readable, and it punishes thin pages that only work when a model is forced to guess. Assume the model is reading everything you publish, not the first paragraph.

Where buyers meet this model

Buyers meet GPT-5.6 Luna through the OpenAI API, usually without knowing it, because it tends to be the model a product routes to for its everyday requests. It also sits underneath OpenAI's consumer chat surfaces as the fast responder, which is where most ordinary questions about vendors and products get answered.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

What is GPT-5.6 Luna used for?
High-volume, latency-sensitive work: chat interfaces, classification and tagging jobs, and lightweight agentic workflows where a task runs many times and each run has to be cheap. It handles text, images and files as input, so document triage and screenshot reading sit within its remit.
How does GPT-5.6 Luna differ from the rest of the GPT-5.6 series?
Luna is the fast, cost-efficient member of the series. It trades the deeper reasoning of its larger siblings for speed and a low per-token cost, which makes it the right pick for volume and the wrong pick for work where a single hard answer has to be correct.
Does GPT-5.6 Luna handle images and documents?
Yes. It accepts text, image and file input, so you can pass it screenshots, scanned pages or uploaded documents alongside a prompt without routing to a separate vision model.
Is GPT-5.6 Luna cheap enough to run on every request?
That is the case it was built for. Output costs several times more than input, so the economics favour reading a lot and writing a little, which is exactly the shape of classification, extraction and routing work.
Surge45°

Is GPT-5.6 Luna recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.