Gemini 2.5 Flash Lite is Google's lightweight multimodal model with a 1M token context window, optimized for high-speed text and image processing.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.050 | $0.200 | $0.010 | |
| $0.100 | $0.400 | $0.010 |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Gemini 2.5 Flash Lite is designed for applications requiring fast, cost-effective AI processing with multimodal capabilities. Its high inference speed and quick response times make it suitable for customer service chatbots, real-time content moderation, rapid document analysis, and interactive applications where latency matters. The large context window enables processing of lengthy documents, code reviews, or multiple images simultaneously, while the lightweight nature keeps operational costs manageable for high-volume deployments. Organizations needing multimodal AI for production applications with tight latency requirements or budget constraints will find this model appropriate for tasks that don't require the full reasoning power of flagship models.
Gemini 2.5 Flash Lite pricing varies by provider and usage type. As a lightweight tier model, it's positioned for cost-efficient deployment. Check the pricing table above for current rates across all available providers.
Gemini 2.5 Flash Lite excels at high-speed multimodal tasks including rapid document processing, real-time chat applications, content moderation, and customer service automation. Its fast 284.5 tokens/second output and 426ms response time make it ideal for latency-sensitive applications requiring both text and image understanding.
As the lightweight variant in the Gemini family, Flash Lite prioritizes speed and efficiency over maximum capability. It maintains the 1M token context window and multimodal support of its siblings while offering faster inference speeds, making it suitable for applications where response time and cost matter more than advanced reasoning capabilities.