Gemini 3.1 Flash Lite is Google's lightweight multimodal model offering fast inference across text, image, audio, and video with a 1M token context window.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.125 | $0.750 | $0.013 | |
| $0.250 | $1.50 | - | |
| $0.250 | $1.50 | $0.025 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Gemini 3.1 Flash Lite is designed for applications requiring fast multimodal processing without the complexity of tool calling or maximum reasoning capability. Its combination of speed, large context window, and broad modality support makes it suitable for content analysis workflows, document processing with mixed media, rapid prototyping of multimodal applications, and high-throughput scenarios where cost efficiency matters. The model works well for summarizing long documents with embedded images, processing video content for basic analysis, and applications needing quick responses across multiple input types. Organizations looking for multimodal capabilities at scale, rather than complex reasoning or agentic workflows, will find this model appropriate for their needs.
Gemini 3.1 Flash Lite pricing varies by provider and usage type. Check the pricing table above for current rates across all available providers and compare input vs output token costs.
Gemini 3.1 Flash Lite excels at fast multimodal processing tasks including document analysis with images, basic video content processing, and high-volume applications where speed and cost efficiency are priorities over maximum reasoning capability.
No, Gemini 3.1 Flash Lite does not include tool calling capabilities. For applications requiring function execution or API integrations, consider Gemini 3.1 Pro or other models that specifically support structured tool calling.