Nova 2.0 Lite is a chat model from Amazon in the Nova family, positioned at the "Lite" tier below the family's larger variants.
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.165 | $1.38 | |
| $0.330 | $2.75 |
Prices updated daily. Last check: Sep 23, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Nova 2.0 Lite suits high-volume text workloads where per-request cost and sustained generation speed matter more than immediate first-token response: batch summarization, content drafting, classification and tagging over large document sets, customer-support reply generation, and back-office text transformation pipelines. The measured ~212 tokens per second output rate means long responses finish quickly once generation starts, which favors asynchronous or queued jobs. The high measured time to first token of about 14 seconds makes it a weaker choice for typing-speed chat interfaces, voice assistants, or autocomplete-style features where users notice startup delay. For tasks that need the deepest reasoning in Amazon's lineup, a larger Nova tier is the more natural comparison point.
Pricing varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are billed differently, and each host sets its own rates. Check the pricing table on this page for current per-provider figures rather than relying on a fixed number.
It is aimed at high-volume, cost-sensitive text generation: summarization, drafting, classification, and batch text processing. Its fast measured output rate of around 212 tokens per second helps on long generations, while its high measured time to first token makes it less suitable for real-time interactive chat.
The "Lite" label places it in the lower-cost tier of the Nova 2.0 family, below the larger variants. In practice that means it is the tier you reach for when throughput and unit cost drive the decision, and you step up to a bigger Nova model when a task needs more capability.
It depends which part of latency you care about. Artificial Analysis measures roughly 212 output tokens per second, which is fast generation, but a time to first token of about 14.2 seconds, which is slow to start. Throughput-bound batch jobs benefit; latency-bound user-facing chat does not.
Our database does not record confirmed modality or tool-calling details for this entry. Check Amazon's official model documentation and your chosen provider's API reference before designing around vision input or function calling.