Gemini 3.7 Flash
Gemini 3.7 Flash is a lightweight-tier model in Google's Gemini family, with a roughly 1M-token context window and text, image, video, and audio input support.
API Pricing
Cheapest on Google Cloud — 49% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.375 | $1.88 | $0.037 | |
| $0.375 | $1.88 | $0.037 | |
| $0.750 | $3.75 | - | |
| $0.750 | $3.75 | $0.075 | |
| $0.750 | $3.75 | $0.075 | |
| $0.750 | $3.75 | $0.075 | |
| $1.35 | $6.75 | - |
Prices updated daily. Last check: Sep 30, 2026
Gemini 3.7 Flash pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond92.1%
- Humanity's Last Exam39.0%
Coding
- SciCode59.8%
Agentic & Tool Use
- Terminal-Bench v2.178.3%
- τ-bench Banking35.5%
Instruction & Long Context
- Long-Context Reasoning83.0%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Family
- Gemini
- Tier
- Lightweight
- Context Window
- 1.0M
- Modalities
- Text, Image, Video, Audio
Capabilities
- Tool Calling
- Yes
- Open Source
- No
- Aliases
- gemini-3.7-flash, Gemini 3.7 Flash, models/gemini-3.7-flash
Strengths & Limitations
Strengths
- Context window of 1,048,576 tokens supports long documents, large codebases, and extended transcripts in a single request
- Accepts image, video, and audio inputs in addition to text, so a single model can handle mixed-media pipelines
- Video and audio input avoid the need for separate transcription or frame-extraction services in many workflows
- Tool calling is supported, enabling function-calling loops and agent integrations
- Flash tier is positioned for high-volume, latency-sensitive workloads rather than only occasional complex queries
- Part of the Gemini family, so prompts and tooling can often be moved between Flash and Pro tiers with limited rework
Limitations
- As a lightweight-tier model, it is not positioned for the hardest reasoning, math, or long-horizon agentic tasks that Pro-tier Gemini models target
- We do not currently track published benchmark scores for this model, so capability comparisons rely on your own testing
- Verified throughput and time-to-first-token measurements are not available in our data
- Very long prompts approaching the 1M-token limit increase latency and token spend regardless of tier
- Provider availability for Gemini models is narrower than for open-weight alternatives, limiting deployment options
Key Features
About Gemini 3.7 Flash
Common Use Cases
Gemini 3.7 Flash suits workloads that combine large inputs with high request volume: summarizing long reports or contracts, extracting structured fields from documents, transcribing and analyzing audio, describing or indexing video content, classification and routing at scale, and customer-facing chat where response latency matters. The million-token context makes it a candidate for whole-repository code question answering or analyzing long meeting recordings without chunking, while tool calling lets it serve as the execution layer in an agent that queries APIs or databases. For tasks that require deep multi-step reasoning, difficult mathematics, or extended autonomous planning, a Pro-tier Gemini model or another reasoning-focused model is generally the better fit, with Flash reserved for the high-frequency steps around it.
Frequently Asked Questions
How much does Gemini 3.7 Flash cost?
Pricing depends on which provider you use and the pricing model — input versus output tokens, batch versus real-time, and any long-context surcharges. Because rates change frequently, check the pricing table on this page for current per-provider figures rather than relying on a fixed number.
What is Gemini 3.7 Flash best used for?
It is aimed at high-volume, latency-sensitive tasks that still need large inputs: document and transcript summarization, structured data extraction, classification and routing, media understanding across image, video, and audio, and tool-calling agent steps. Save the hardest reasoning tasks for a Pro-tier model.
How does it differ from Gemini Pro-tier models?
Flash is Google's lightweight tier, tuned toward throughput and cost efficiency, while Pro-tier Gemini models target deeper reasoning and more complex agentic work. Both share the Gemini family's long-context and multimodal design, so many prompts port between them, but you should expect Pro tiers to handle harder problems and Flash to handle more requests per unit of spend.
Can Gemini 3.7 Flash process video and audio files?
Yes — our metadata confirms text, image, video, and audio inputs. That allows a single API call to reason over a video clip or audio recording alongside written instructions, which can remove separate transcription or frame-sampling steps from a pipeline.
Does it support function calling for agents?
Tool calling is supported, so the model can be connected to external functions, APIs, and databases within agent frameworks. Specific parameters such as parallel calls or strict schema enforcement depend on the provider's API implementation, so check the provider documentation you plan to use.
How large is the context window in practice?
The context window is 1,048,576 tokens, roughly one million. That is enough for very long documents, multi-hour transcripts, or large portions of a codebase in one prompt, though filling the window increases both latency and token consumption, so most production deployments still trim inputs to what is needed.