Gemini 2.0 Flash-Lite is a lightweight chat model from Google, positioned as the smaller, higher-throughput option within the Gemini 2.0 Flash line.
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.037 | $0.150 | |
| $0.075 | $0.300 |
Prices updated daily. Last check: Sep 22, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Gemini 2.0 Flash-Lite is aimed at workloads that run constantly and need predictable, low per-request cost: intent and topic classification, sentiment tagging, structured field extraction from documents or emails, short summarization, content moderation triage, query rewriting for search and RAG pipelines, and first-pass drafting inside product UIs. It also works well as the cheap tier in a routing setup, where it handles the bulk of straightforward traffic and hands ambiguous or complex requests up to Gemini 2.0 Flash or a Pro-tier model. It is a weaker fit for long agentic chains, hard mathematical or scientific reasoning, and large-scale code generation, where a higher-tier model in the Gemini family or a dedicated reasoning model is the more appropriate choice.
Pricing depends on the provider and on the pricing type — input tokens, output tokens, cached input, and batch requests are typically billed at different rates, and resellers may price differently from Google's own API. Check the pricing table on this page for current per-provider rates rather than relying on any figure quoted elsewhere.
High-volume, well-defined text tasks: classification, extraction, short summarization, query rewriting, moderation triage, and as the low-cost first tier in a model-routing pipeline. It is less suited to deep multi-step reasoning or complex code generation.
Both are in the Gemini 2.0 Flash line, but Flash-Lite is the lighter-weight variant positioned below standard Flash. Flash-Lite is chosen when cost per request and speed dominate; standard Flash offers more capability headroom for harder prompts. Many teams run both and route between them.
Our database does not carry verified modality or context-window fields for this entry, so we do not state either way here. Consult Google's official model documentation for the authoritative capability list before building against those features.
Gemini models are offered as a hosted API service by Google rather than as downloadable weights. If your requirement is to run inference on rented or owned hardware, compare open-weight alternatives and see our GPU pricing pages for the compute side of that cost.