See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
NVIDIA model°

Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's open frontier-reasoning and orchestration model, the largest entry in the Nemotron 3 line, released under an open weights licence for teams who want a reasoning model they can run themselves.

NVIDIAVerified 04/10/2026Released 04/06/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 232 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.50

field median $0.41

Output / 1M tokens

$2.20

field median $2.00

Context

262K

field median 500K

02 / overview

What Nemotron 3 Ultra is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Nemotron 3 Ultra is a text-only reasoning and orchestration model, meaning it is built to work through multi-step problems and to direct other models and tools rather than to hold a light conversation. NVIDIA describes it as a mixture-of-experts model where only a fraction of its total parameters are active on any given token, which is what keeps a model of this size affordable to serve. It does not handle images, audio or video, so anything involving documents as pictures or screenshots needs a different model in front of it.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Nemotron 3 Ultra first appeared in June 2026 as the top tier of NVIDIA's Nemotron 3 family. Reach for it when the work is genuinely hard reasoning or when one model needs to plan and route across a set of other models and tools, and the weights being open matters to you. It is the wrong pick for high-volume, simple classification or chat, where a smaller Nemotron or a lighter hosted model will do the same job for less.

How you reach it

The API, the apps it powers, and what its limits let you do.

Nemotron 3 Ultra is available through NVIDIA's own API endpoints and, because the weights are open, through third-party inference hosts and self-hosted deployments on your own GPUs. The context window runs to a quarter of a million tokens and the output ceiling is unusually generous, so it can take in a large codebase or a long document set and return a full reasoning trace, a long plan or an extended piece of generated code in a single call rather than being chunked. Everything in and out is plain text.

Why it matters

What changes because this exists, or why it does not.

An open model at this level of reasoning changes where the work can live. Teams that could not send data to a hosted frontier API, or that wanted to fine-tune a strong reasoner rather than prompt one, now have a serious option, and the very long output limit means long-form agent plans and large generated artefacts no longer need to be stitched together across calls.

Follows Nemotron 3.5 Content Safety. Superseded by Nemotron 3.5 Lightning. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Nemotron 3 Ultra API pricingSurge45°
ChargePriceUnitRead onSource
Input$0.50per 1M tokens2026-10-03Check
Output$2.20per 1M tokens2026-10-03Check

For a model in the frontier-reasoning bracket, Nemotron 3 Ultra is priced toward the cheap end, with output costing four times input in the usual pattern. The sparse mixture-of-experts design is what makes that possible, and because the weights are open you can also take the self-hosting route and trade the per-token bill for your own GPU capacity.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$21.00
A busy support assistant200M tokens40M tokens$188.00
A document pipeline1000M tokens100M tokens$720.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

Price history (6 earlier readings)
  • 2026-10-02: Output $2.20 per 1M tokens
  • 2026-10-02: Input $0.60 per 1M tokens
  • 2026-10-02: Output $2.40 per 1M tokens
  • 2026-10-02: Input $0.50 per 1M tokens
  • 2026-09-24: Output $2.40 per 1M tokens
  • 2026-09-24: Input $0.60 per 1M tokens

Every reading we have taken, kept whether it changed or not. This is the part nobody can copy in a week.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window262,144 tokens
Maximum output16,384 tokens
Modalitiestext
Released04/06/2026
StatusCurrent
Catalogue identifiernvidia/nemotron-3-ultra-550b-a55b

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from NVIDIA

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Nemotron 3 Super11/03/2026
  2. 02Nemotron 3 Nano Omni (free)28/04/2026
  3. 03Nemotron 3.5 Content Safety04/06/2026
  4. 04Nemotron 3 Ultra04/06/2026
  5. 05Nemotron 3.5 Lightning11/08/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Nemotron 3 Ultra will not cite your brand to an end user directly, but it is the kind of model that sits underneath agents and internal tools that do the retrieving and summarising. The practical implication is the same as ever: your content needs to be clean, well-structured text that a long-context reasoner can pull in wholesale and reason over, because a model reading a quarter of a million tokens of your documentation will form a view of your product from whatever is there.

Where buyers meet this model

There is no consumer chat app behind Nemotron 3 Ultra, so buyers do not meet it by name the way they meet a mainstream assistant. They encounter it through NVIDIA's API, through inference providers hosting the open weights, and, most often, without knowing it, as the reasoning or orchestration layer inside a vendor's own product or agent.

Change log

  • Nemotron 3 Ultra pricing changed

    Model: Nemotron 3 Ultra Input was $0.6 per 1M tokens, now $0.5 per 1M tokens Output was $2.4 per 1M tokens, now $2.2 per 1M tokens

    Pricing

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Nemotron 3 Ultra open source?
NVIDIA describes it as an open model, so the weights are available to download, run and fine-tune rather than being reachable only through a hosted API. Check the specific licence terms before commercial deployment.
Can Nemotron 3 Ultra read images or PDFs?
No. It is text-only, so any visual input needs to be converted to text by another tool before it reaches the model.
What does orchestration mean here?
NVIDIA positions Nemotron 3 Ultra as a model that plans and directs work across other models and tools, rather than one that only answers prompts itself. In practice that means it sits at the centre of an agent system deciding what happens next.
Should we use it instead of a hosted frontier model?
It is worth considering if you need open weights for data residency, fine-tuning or cost control on high volumes. If you want a multimodal model or a consumer-facing assistant, look elsewhere.
Surge45°

Is Nemotron 3 Ultra recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.