Nemotron 3 Ultra 550B A55B is a large-scale model from NVIDIA in its Nemotron 3 family, positioned at the Ultra tier with roughly 550B total and 55B active parameters per token.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.600 | $3.60 | - | |
| $0.600 | $2.40 | $0.120 | |
| $0.800 | $2.60 | $0.100 | |
| $1.00 | $3.20 | $0.250 |
Prices updated daily. Last check: Sep 22, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Nemotron 3 Ultra 550B A55B is aimed at the workloads where teams reach for the largest model in a family: multi-step reasoning chains, code generation and refactoring across larger contexts, document analysis, and agentic systems where an error early in a chain compounds. Its measured throughput of roughly 139 tokens per second makes it usable for streaming assistant interfaces and background batch jobs alike, though the ~1.3 second time to first token makes it a poor fit for ultra-low-latency tasks like inline autocomplete or high-volume classification — those are better served by smaller Nemotron 3 tiers or other lightweight models. Because the Ultra tier represents the top of NVIDIA's Nemotron 3 range, a common pattern is to route only the hardest requests here and handle the rest with cheaper models, comparing per-provider costs in the pricing table on this page to decide where that cutoff should fall.
Pricing varies by provider and by pricing type — hosted token-based APIs, batch or cached-input rates, and self-managed GPU deployments all price differently, and rates change frequently. Check the pricing table on this page for current per-provider figures rather than relying on a fixed number.
It is best suited to the harder end of a workload mix: multi-step reasoning, code generation and review, document analysis, synthetic data generation, and agentic pipelines where accuracy at each step matters. Lower-complexity, high-volume traffic is usually more economical on smaller Nemotron 3 tiers.
It follows NVIDIA's naming convention for sparse mixture-of-experts models: roughly 550 billion total parameters with about 55 billion active for any given token. The practical effect is that inference compute scales with the active parameter count, not the full total, while the model still draws on the larger parameter pool.
Artificial Analysis measured approximately 139 output tokens per second with a time to first token of about 1347 ms. Actual figures depend heavily on the provider, hardware, batch size, and prompt length, so treat these as a reference point rather than a guarantee.
We do not have a confirmed context window recorded for Nemotron 3 Ultra 550B A55B. Check NVIDIA's model card and your chosen provider's documentation before designing around long-context inputs, since hosts sometimes serve a shorter window than the model supports.
The Ultra designation places it at the top of the Nemotron 3 sizing range, so it is intended for tasks that smaller family members struggle with. The usual trade-off applies: more capacity per request, but higher cost per token and slower first-token latency than the smaller tiers.