Nemotron 3.5 Lightning
Nemotron 3.5 Lightning is a model in NVIDIA's Nemotron family, positioned as a speed-oriented variant with measured output throughput around 291 tokens per second.
API Pricing
Cheapest on Crusoe — 25% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.050 | $0.200 | $0.030 | |
| $0.071 | $0.195 | $0.036 | |
| $0.080 | $0.200 | $0.040 |
Prices updated daily. Last check: Sep 26, 2026
Nemotron 3.5 Lightning pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond74.3%
- Humanity's Last Exam10.6%
Coding
- SciCode32.1%
Agentic & Tool Use
- Terminal-Bench v2.124.3%
- τ-bench Banking8.9%
Instruction & Long Context
- Long-Context Reasoning60.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- NVIDIA
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Measured output throughput of approximately 291 tokens per second, placing it in the faster tier of served models tracked by Artificial Analysis
- Time to first token around 782 ms, keeping initial response latency under one second in benchmark conditions
- Speed profile suits streaming chat UIs where perceived responsiveness matters
- Part of NVIDIA's Nemotron family, so it sits alongside sibling variants that can be swapped in for different speed/quality points
- High token rate reduces wall-clock time on long-form generation and multi-step agent loops
- Published third-party performance measurements available rather than vendor-only claims
Limitations
- Speed-oriented variants generally trade some depth on hard reasoning tasks relative to larger siblings
- We do not track a confirmed context window figure for this model — verify limits with the specific provider
- Modality support (for example image input) is not confirmed in our metadata
- Quality benchmark scores are not part of our verified data for this model, so capability comparisons must come from other sources
- Serving availability and configuration vary between providers, so measured latency may differ from the benchmark figure
Key Features
About Nemotron 3.5 Lightning
Common Use Cases
Nemotron 3.5 Lightning fits workloads where throughput and responsiveness are the binding constraints: interactive chat and assistant front-ends, streaming completions where users watch text appear, summarization and rewriting over large document batches, and agent or RAG pipelines that issue many sequential model calls per user request. Its sub-second time to first token makes it a reasonable candidate for voice or real-time interfaces where the initial pause is noticeable, and its ~291 tokens per second output rate shortens wall-clock time on long generations. For tasks that hinge on multi-step mathematical or scientific reasoning, teams commonly route those requests to a heavier reasoning model and reserve a Lightning-class model for the high-volume path.
Frequently Asked Questions
How much does Nemotron 3.5 Lightning cost to use?
Pricing depends on which provider hosts the model and on the pricing type — separate input and output token rates are typical, and some providers offer batch or dedicated-capacity options. Rates change frequently, so see the pricing table on this page for current per-provider figures.
What is Nemotron 3.5 Lightning best used for?
It is best suited to high-volume, latency-sensitive workloads: interactive chat, streaming generation, batch summarization and rewriting, and multi-call agent or retrieval pipelines. Its measured ~291 tokens per second output rate and ~782 ms time to first token are its main documented advantages.
How fast is Nemotron 3.5 Lightning?
Artificial Analysis measured roughly 291 output tokens per second with a time to first token of about 782 milliseconds. Actual figures vary by provider, region, prompt length and load, so treat the benchmark as a reference point rather than a guarantee.
Who makes Nemotron 3.5 Lightning?
NVIDIA develops the Nemotron model line, of which Nemotron 3.5 Lightning is a member. It is served through third-party inference providers, which are listed with their rates in the pricing table above.
What context window does Nemotron 3.5 Lightning support?
We do not have a confirmed context window figure for this model in our database. Providers sometimes expose different maximum context lengths for the same model, so check the documentation of the specific provider you plan to use.