Skip to main content
NVIDIA

Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is a model in NVIDIA's Nemotron family, positioned as a speed-oriented variant with measured output throughput around 291 tokens per second.

Input from
$0.050 / 1M tokens
across 3 providers

API Pricing

Cheapest on Crusoe — 25% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.050$0.200$0.030
$0.071$0.195$0.036
$0.080$0.200$0.040

Prices updated daily. Last check: Sep 26, 2026

Nemotron 3.5 Lightning pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
12.9 / 100
Coding
26.8 / 100
Output Speed
273 t/s
Latency (TTFT)
452ms

Reasoning & Knowledge

  • GPQA Diamond74.3%
  • Humanity's Last Exam10.6%

Coding

  • SciCode32.1%

Agentic & Tool Use

  • Terminal-Bench v2.124.3%
  • τ-bench Banking8.9%

Instruction & Long Context

  • Long-Context Reasoning60.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
NVIDIA
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Measured output throughput of approximately 291 tokens per second, placing it in the faster tier of served models tracked by Artificial Analysis
  • Time to first token around 782 ms, keeping initial response latency under one second in benchmark conditions
  • Speed profile suits streaming chat UIs where perceived responsiveness matters
  • Part of NVIDIA's Nemotron family, so it sits alongside sibling variants that can be swapped in for different speed/quality points
  • High token rate reduces wall-clock time on long-form generation and multi-step agent loops
  • Published third-party performance measurements available rather than vendor-only claims

Limitations

  • Speed-oriented variants generally trade some depth on hard reasoning tasks relative to larger siblings
  • We do not track a confirmed context window figure for this model — verify limits with the specific provider
  • Modality support (for example image input) is not confirmed in our metadata
  • Quality benchmark scores are not part of our verified data for this model, so capability comparisons must come from other sources
  • Serving availability and configuration vary between providers, so measured latency may differ from the benchmark figure

Key Features

•Output generation measured at ~291 tokens per second (Artificial Analysis)
•Time to first token measured at ~782 ms
•Latency-oriented "Lightning" variant within the Nemotron 3.5 generation
•Developed by NVIDIA as part of the Nemotron model line
•Available through inference API providers listed in the pricing table
•Streaming-friendly performance profile for interactive applications

About Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is a large language model from NVIDIA, part of the company's Nemotron line of models. The "Lightning" designation places it as a latency- and throughput-oriented member of the Nemotron 3.5 generation, aimed at workloads where response speed and serving cost per request matter as much as raw answer quality. Independent measurements collected by Artificial Analysis put the model's output speed at roughly 291 tokens per second with a time to first token of about 782 milliseconds. That combination — high sustained generation rate paired with a sub-second initial latency — is the model's most clearly documented characteristic in the data we track. Other specifications for this model, including context window length and supported input modalities, are not part of our verified metadata, so we do not report them here; check the provider listings in the pricing table for the exact configuration each host exposes. In practice, models positioned this way in a family are typically used for high-volume generation, chat interfaces where perceived responsiveness drives user experience, and pipeline stages that run many calls per task. Buyers comparing Nemotron 3.5 Lightning against other options generally weigh its measured throughput against heavier reasoning-focused models that trade tokens per second for depth on hard problems.

Common Use Cases

Nemotron 3.5 Lightning fits workloads where throughput and responsiveness are the binding constraints: interactive chat and assistant front-ends, streaming completions where users watch text appear, summarization and rewriting over large document batches, and agent or RAG pipelines that issue many sequential model calls per user request. Its sub-second time to first token makes it a reasonable candidate for voice or real-time interfaces where the initial pause is noticeable, and its ~291 tokens per second output rate shortens wall-clock time on long generations. For tasks that hinge on multi-step mathematical or scientific reasoning, teams commonly route those requests to a heavier reasoning model and reserve a Lightning-class model for the high-volume path.

Frequently Asked Questions

How much does Nemotron 3.5 Lightning cost to use?

Pricing depends on which provider hosts the model and on the pricing type — separate input and output token rates are typical, and some providers offer batch or dedicated-capacity options. Rates change frequently, so see the pricing table on this page for current per-provider figures.

What is Nemotron 3.5 Lightning best used for?

It is best suited to high-volume, latency-sensitive workloads: interactive chat, streaming generation, batch summarization and rewriting, and multi-call agent or retrieval pipelines. Its measured ~291 tokens per second output rate and ~782 ms time to first token are its main documented advantages.

How fast is Nemotron 3.5 Lightning?

Artificial Analysis measured roughly 291 output tokens per second with a time to first token of about 782 milliseconds. Actual figures vary by provider, region, prompt length and load, so treat the benchmark as a reference point rather than a guarantee.

Who makes Nemotron 3.5 Lightning?

NVIDIA develops the Nemotron model line, of which Nemotron 3.5 Lightning is a member. It is served through third-party inference providers, which are listed with their rates in the pricing table above.

What context window does Nemotron 3.5 Lightning support?

We do not have a confirmed context window figure for this model in our database. Providers sometimes expose different maximum context lengths for the same model, so check the documentation of the specific provider you plan to use.