See how ChatGPT, Perplexity and Google AI Overviews describe you today

Free AI Visibility Audit
NVIDIA model°

Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is an open, text-only mixture-of-experts model from NVIDIA, built so that only a small fraction of its total parameters are active on any given request. It is aimed at high-throughput agentic work, the kind where a system makes thousands of small model calls rather than a few large ones.

NVIDIAVerified 04/10/2026Released 11/08/2026

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 232 models we track, not against an absolute standard.

Capability

Not measured

No published benchmark scores yet

Input / 1M tokens

$0.06

field median $0.41

Output / 1M tokens

$0.17

field median $2.00

Context

262K

field median 500K

02 / overview

What Nemotron 3.5 Lightning is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

What it is

The job it was built for, and the job it is not for.

Nemotron 3.5 Lightning is a small-active-parameter mixture-of-experts model released openly by NVIDIA, handling text in and text out. The job it was built for is volume: agent loops, routing steps, classification, extraction and other specialised tasks that need to run cheaply and often. It is not a multimodal model and not the one to reach for when a single answer needs to be as good as it can possibly be.

When it arrived, and when to use it

Where it sits in its line, and when a sibling is the better pick.

Nemotron 3.5 Lightning appeared in August 2026 as the lightweight member of NVIDIA's Nemotron 3.5 line. It is the right pick when the same prompt pattern runs at scale and latency and unit cost decide whether the feature ships at all. It is the wrong pick for one-off reasoning of the sort a customer reads directly, or for anything involving images, audio or long-form authored prose.

How you reach it

The API, the apps it powers, and what its limits let you do.

Being an open model, Nemotron 3.5 Lightning can be pulled down and served on your own NVIDIA hardware, or reached through hosted inference providers that carry it. It takes text only, with a context window large enough to hold a substantial document set or a long agent trace in a single call, and an output ceiling sized for structured results rather than book-length generations. In practice that shape suits tool-calling agents, retrieval pipelines and batch jobs over large bodies of text.

Why it matters

What changes because this exists, or why it does not.

Cheap models existed before Nemotron 3.5 Lightning, but few combined an open licence, a context window this wide and an active-parameter count this low. That makes it realistic to run an agent that reasons across a whole repository or ticket history on every step, instead of trimming context to keep the bill down. Because it is open, teams with their own GPUs can take the per-call cost close to their electricity bill.

Follows Nemotron 3 Ultra, and is the newest in its line. See the whole line.

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means we read it on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

No published benchmark scores for this model yet.

We publish a score only where we can link the result to where it was published. Until a benchmark result for this model exists in a source we read, this section stays empty rather than being filled with a provider's marketing figure.

How we decide what counts as evidence

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Nothing in the ledger yet.

Each row here carries a score, the benchmark it came from, what the leading model scored on the same test, and a link to the published result. Rows appear as results are published and read.

How we decide what counts as evidence

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against everything we track: ranking against models nobody tested would rank who published, not who is better.

No category scores to shape yet.

A category score is the weighted mean of the benchmarks published for it. With no published rows there is nothing to average, and an empty chart drawn at zero would say something false.

How we decide what counts as evidence

06 / cost

What it costs

List API rates as last read from the provider, with the source on every row, plus every change we have recorded since we started tracking it.

Nemotron 3.5 Lightning API pricingSurge45°
ChargePriceUnitRead onSource
Output$0.17per 1M tokens2026-10-03Check
Input$0.06per 1M tokens2026-10-03Check

Nemotron 3.5 Lightning sits at the very bottom of the market on both input and output, in the bracket where token cost stops being the thing that shapes your architecture. Self-hosting removes the per-token charge entirely and replaces it with the cost of the hardware you already run, which is much of the point of an open release.

What a month costsSurge45°
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$2.04
A busy support assistant200M tokens40M tokens$18.70
A document pipeline1000M tokens100M tokens$76.50

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

Price history (12 earlier readings)
  • 2026-10-02: Input $0.06 per 1M tokens
  • 2026-10-02: Output $0.16 per 1M tokens
  • 2026-10-01: Output $0.17 per 1M tokens
  • 2026-10-01: Input $0.06 per 1M tokens
  • 2026-09-29: Input $0.06 per 1M tokens
  • 2026-09-29: Output $0.16 per 1M tokens
  • 2026-09-26: Input $0.08 per 1M tokens
  • 2026-09-26: Output $0.20 per 1M tokens
  • 2026-09-25: Input $0.07 per 1M tokens
  • 2026-09-25: Output $0.20 per 1M tokens
  • 2026-09-24: Output $0.20 per 1M tokens
  • 2026-09-24: Input $0.08 per 1M tokens

Every reading we have taken, kept whether it changed or not. This is the part nobody can copy in a week.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationSurge45°
Context window262,144 tokens
Maximum output131,072 tokens
Modalitiestext
Released11/08/2026
StatusCurrent
Catalogue identifiernvidia/nemotron-3.5-lightning

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from NVIDIA

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Came after

Nothing newer in this line yet.

  1. 01Nemotron 3 Super11/03/2026
  2. 02Nemotron 3 Nano Omni (free)28/04/2026
  3. 03Nemotron 3.5 Content Safety04/06/2026
  4. 04Nemotron 3 Ultra04/06/2026
  5. 05Nemotron 3.5 Lightning11/08/2026

Ordered by release date and worked out from the naming, so a new member slots in as soon as its page exists. A retired model keeps its page and its place in the line.

10 / notes

Our notes

What this model changes for a brand trying to be cited in AI answers, and every change we have logged since it launched.

What it changes for you

Models like Nemotron 3.5 Lightning make it cheap enough to read everything, so more of the pipelines that decide which brands get cited will ingest your full documentation rather than a snippet. That rewards pages that are clean, self-contained and easy to parse without a human in the loop. It also means the model doing the reading may be a small open one running on someone else's infrastructure, not a frontier model you can test against directly.

Where buyers meet this model

There is no consumer chat app for Nemotron 3.5 Lightning, so buyers almost never meet it head on. They meet it underneath things: the retrieval and classification layers of AI search and assistant products, agent frameworks that fan out many small calls, and internal tools built by teams running the weights themselves. If your content is being read, summarised or scored by a pipeline rather than a chatbot, a model of this class is often the one doing it.

Change log

  • Nemotron 3.5 Lightning cuts input and output token prices

    NVIDIA has lowered pricing for Nemotron 3.5 Lightning on OpenRouter. Input drops from $0.08 to $0.06 per 1M tokens, and output drops from $0.20 to $0.16 per 1M tokens. That is a 25 percent cut on input and 20 percent on output.

    Pricing
  • Nemotron 3.5 Lightning pricing changed

    Model: Nemotron 3.5 Lightning Input was $0.07 per 1M tokens, now $0.08 per 1M tokens Output was $0.2 per 1M tokens, now $0.2 per 1M tokens

    Pricing
  • Nemotron 3.5 Lightning pricing changed

    Model: Nemotron 3.5 Lightning Input was $0.08 per 1M tokens, now $0.07 per 1M tokens Output was $0.2 per 1M tokens, now $0.2 per 1M tokens

    Pricing

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

Is Nemotron 3.5 Lightning open?
Yes. NVIDIA released it as an open mixture-of-experts model, so it can be self-hosted rather than only accessed through a vendor API.
Can Nemotron 3.5 Lightning handle images?
No. It is a text-only model, so anything involving images, audio or video needs a different model in the pipeline.
What is Nemotron 3.5 Lightning best used for?
High-throughput agentic workloads and specialised repeated tasks, such as tool-calling loops, routing, extraction and classification, where cost and latency per call matter more than peak single-answer quality.
Surge45°

Is Nemotron 3.5 Lightning recommending you?

Models change what gets cited. We measure whether AI answers name your brand or your competitors across every assistant, and show you what to change.