Gemma 4 31B is Google's lightweight multimodal model supporting text, image, and video inputs with a 262K token context window.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.090 | $0.340 | $0.050 | |
| $0.090 | $0.340 | $0.050 | |
| $0.140 | $0.400 | - | |
| $0.390 | $0.970 | - | |
| $0.390 | $0.970 | - |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Gemma 4 31B is well-suited for applications requiring multimodal analysis where efficiency is important, such as content moderation across text and visual media, educational tools that process mixed-media content, and customer support systems handling images or videos alongside text queries. The model's lightweight nature makes it appropriate for high-volume scenarios like document analysis with embedded images, social media content processing, or applications where consistent low latency is preferred over maximum reasoning capability. Its video input support enables use cases like surveillance analysis, educational video summarization, and media content categorization where real-time or high-throughput processing is valued.
Gemma 4 31B pricing varies by provider and may differ for text versus multimodal inputs. Check the pricing table above for current rates across all providers offering this model.
Gemma 4 31B excels at multimodal tasks requiring text, image, and video understanding where efficiency is important. It's ideal for content analysis, document processing with visuals, media moderation, and applications needing substantial context length without the computational cost of flagship models.
No, Gemma 4 31B does not support tool calling or function calling capabilities. For applications requiring API integrations or external tool use, you would need to consider other models that include these features.