See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
OpenAI model°

GPT-4.1 Mini

GPT-4.1 Mini is OpenAI's mid-sized model in the GPT-4.1 family, built to match GPT-4o-class quality at lower latency and cost while keeping a million-token context window.

OpenAIVerified 24/09/2026Released 14/04/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 126 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.40

field median $0.30

Output / 1M tokens

$1.60

field median $1.25

Context

1,048K

field median 524K

02 / overview

What GPT-4.1 Mini is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

GPT-4.1 Mini is a general-purpose text model that also accepts images and files, aimed at work that needs solid reasoning over long inputs without paying for a frontier tier. It is the model you reach for when volume matters: classification, extraction, summarising long documents, powering an assistant that runs constantly. It is not positioned as OpenAI's strongest reasoning model, so the hardest multi-step problems belong elsewhere.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

GPT-4.1 Mini was first seen in April 2025, arriving alongside the rest of the GPT-4.1 line as the middle option between the full model and the smallest one. It is the right pick when you want most of the capability of the larger GPT-4.1 at a fraction of the running cost, particularly on high-request workloads. It is the wrong pick when a task genuinely needs the top of the family, or when latency and cost matter so much that the smallest sibling would do the job.

How you reach it

The API, the apps it powers, and what its limits let you do.

GPT-4.1 Mini is reached through the OpenAI API and is available in the tooling built on it, so most teams meet it as a model string rather than a separate product. The million-token context window means you can put entire codebases, document sets or long transcripts into a single call instead of building a retrieval layer first, and image and file input let it read screenshots, scans and PDFs in the same request. Output length is capped well below the input window, so it is better at reading a lot and answering briefly than at generating very long documents.

Why it matters

What changes because this exists, or why it does not.

The practical change is that long-context work stopped being a premium decision. Feeding a whole quarter of support tickets or a full documentation set into one prompt is now cheap enough to do on every request rather than as a batch job, which removes a lot of chunking and retrieval plumbing from ordinary pipelines.

Follows GPT-4.1. Superseded by GPT-4.1 Nano. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

GPT-4.1 Mini API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.40per 1M tokens2026-09-24Check
Output$1.60per 1M tokens2026-09-24Check

Running GPT-4.1 Mini costs a small fraction of a frontier model, with output priced at several times input as is standard across the field. That puts it in the bracket where you can afford to send long documents on every call, which is the main reason teams pick it over the larger models in the same family.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$16.00
A busy support assistant200M tokens40M tokens$144.00
A document pipeline1000M tokens100M tokens$560.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window1,047,576 tokens
Maximum output32,768 tokens
Modalitiesimage, text, file
Released14/04/2025
StatusCurrent
Catalogue identifieropenai/gpt-4.1-mini

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from OpenAI

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came before

GPT-4.1
  1. 01GPT-4.114/04/2025
  2. 02GPT-4.1 Mini14/04/2025
  3. 03GPT-4.1 Nano14/04/2025
  4. 04GPT-507/08/2025
  5. 05GPT-5 Mini07/08/2025
  6. 06GPT-5 Nano07/08/2025
  7. 07GPT-5 Pro06/10/2025
  8. 08GPT-5.113/11/2025
  9. 09GPT-5.1-Codex13/11/2025
  10. 10GPT-5.1-Codex-Max04/12/2025
  11. 11GPT-5.210/12/2025
  12. 12GPT-5.2 Pro10/12/2025
  13. 13GPT-6 Astra04/09/2026
  14. 14GPT-6 Astra Pro04/09/2026
  15. 15GPT-6 Luna22/09/2026
  16. 16GPT-6 Luna Pro22/09/2026
  17. 17GPT-6 Sol22/09/2026
  18. 18GPT-6 Sol Pro22/09/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Cheap long context means the assistants and search features citing your brand are increasingly reading whole source documents rather than short retrieved snippets. Your documentation, pricing pages and comparison content are more likely to be ingested in full, so gaps and contradictions across pages become visible in a way that snippet-level retrieval used to hide. Write for the reader who has all of it in front of them at once.

Where buyers meet this model

Most buyers meet GPT-4.1 Mini indirectly, as the model quietly running inside a SaaS product's AI features, where the vendor chose it for cost rather than badging it. Developers meet it directly in the OpenAI API when picking a default model for production traffic. It is not a consumer-facing brand in its own right, so nobody arrives asking for it by name.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is GPT-4.1 Mini good enough to replace GPT-4o in production?
OpenAI describes it as competitive with GPT-4o at substantially lower latency and cost, so for many production workloads it is a straight swap. Test it on your own evaluation set first, particularly anything involving multi-step reasoning, since "competitive" is not the same as identical.
What can you actually do with a million-token context window?
You can put an entire codebase, a full document set or a long set of transcripts into a single call rather than building retrieval infrastructure to select passages first. The output cap is much smaller than the input window, so it suits reading a lot and answering concisely rather than producing very long documents.
Does GPT-4.1 Mini handle images and files?
Yes, it takes image, text and file input, so it can work from screenshots, scans and PDFs in the same request as ordinary text. Output is text only.
When should I use the larger GPT-4.1 instead?
When the task genuinely needs the extra capability and the volume is low enough that cost per call is not the constraint. For high-request workloads like classification, extraction and summarisation, GPT-4.1 Mini is usually the better economics.
Surge45°

Is GPT-4.1 Mini recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.