See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
NVIDIA model°

Nemotron 3 Nano Omni (free)

Nemotron 3 Nano Omni is NVIDIA's open multimodal model, a 30B-A3B system built to act as a perception and context sub-agent inside larger enterprise agent stacks, and it is offered free to run.

NVIDIAVerified 04/10/2026Released 28/04/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 232 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

Free

field median $0.41

Output / 1M tokens

Free

field median $2.00

Context

256K

field median 500K

02 / overview

What Nemotron 3 Nano Omni (free) is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Nemotron 3 Nano Omni is a small open multimodal model that reads text, images, video and audio and turns them into something the rest of an agent system can act on. NVIDIA positions it as a sub-agent, the component that handles perception and context rather than the one that owns the reasoning or the final answer. It is not intended to be the single model behind a general assistant product.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Nemotron 3 Nano Omni appeared in April 2026 as the Nano tier of NVIDIA's Nemotron 3 family, the small end rather than the flagship. It is the right pick when you need many cheap perception calls across mixed media inside a pipeline that has a stronger model above it. It is the wrong pick when you want one model to do the reasoning, the writing and the final customer-facing response on its own.

How you reach it

The API, the apps it powers, and what its limits let you do.

Nemotron 3 Nano Omni is reached through the API as an open model, which means it can also be pulled and self-hosted rather than only rented. The wide context window lets you hand it long transcripts, document sets or extended video and audio alongside text in a single call, and the large output ceiling means it can return full structured extractions rather than short summaries. In practice it sits inside agent frameworks, feeding parsed context to whatever model does the deciding.

Why it matters

What changes because this exists, or why it does not.

A four-modality model at this size, free to call and open to host, makes it reasonable to run perception on everything instead of triaging what is worth processing. Work that previously meant stitching together a separate speech model, a vision model and a document parser can collapse into one component. That is a plumbing change more than a capability leap, but it is the kind of change that shows up in monthly bills.

Follows Nemotron 3 Super. Superseded by Nemotron 3.5 Content Safety. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Nemotron 3 Nano Omni (free) API pricingSurge45°
ChargePriceUnitRead onSource
InputFreeper 1M tokens2026-09-24Check
OutputFreeper 1M tokens2026-09-24Check

Nemotron 3 Nano Omni is listed at no cost for both input and output, which puts it at the floor of the market rather than merely cheap. The real cost sits in the infrastructure if you self-host it, so compare it on GPU time and engineering effort, not on token rates.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokensFree
A busy support assistant200M tokens40M tokensFree
A document pipeline1000M tokens100M tokensFree

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window256,000 tokens
Maximum output65,536 tokens
Modalitiestext, audio, image, video
Released28/04/2026
StatusCurrent
Catalogue identifiernvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from NVIDIA

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

  1. 01Nemotron 3 Super11/03/2026
  2. 02Nemotron 3 Nano Omni (free)28/04/2026
  3. 03Nemotron 3.5 Content Safety04/06/2026
  4. 04Nemotron 3 Ultra04/06/2026
  5. 05Nemotron 3.5 Lightning11/08/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Models like this sit upstream of the answer, not at the point of citation, so its arrival does not change who gets named in an AI response. What it does change is the volume and variety of source material an agent can afford to ingest, including video, audio and screenshots of your product. Documentation, demos and recorded walkthroughs become machine-readable inputs in a way they were not when text was the only cheap modality.

Where buyers meet this model

Buyers rarely meet Nemotron 3 Nano Omni directly, there is no consumer chat app attached to it. They encounter it through the API, through NVIDIA's enterprise agent tooling, and through whatever internal product a vendor has built on top of it, where it does the seeing and hearing and something else does the talking.

Change log

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Nemotron 3 Nano Omni really free?
The listed input and output token rates are zero, and the model is open, so you can also run it on your own hardware. Self-hosting still costs you compute and engineering time, so treat it as free to license rather than free to operate.
What can Nemotron 3 Nano Omni actually handle as input?
It accepts text, images, video and audio, which is the full set of modalities most enterprise pipelines need in one component. Output is text.
Should Nemotron 3 Nano Omni be the main model in my product?
NVIDIA describes it as a perception and context sub-agent, which means it is designed to feed a stronger model rather than replace one. Use it for ingestion and extraction, and keep a larger model for reasoning and customer-facing output.
Surge45°

Is Nemotron 3 Nano Omni (free) recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.