Skip to main content
LightweightGoogle

Gemini 3.7 Flash

Gemini 3.7 Flash is a lightweight-tier model in Google's Gemini family, with a roughly 1M-token context window and text, image, video, and audio input support.

Context 1.0M
Tier Lightweight
Tools Supported
Modalities text, image, video, audio
Input from
$0.375 / 1M tokens
across 5 providers

API Pricing

Cheapest on Google Cloud — 49% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.375$1.88$0.037
$0.375$1.88$0.037
$0.750$3.75-
$0.750$3.75$0.075
$0.750$3.75$0.075
$0.750$3.75$0.075
$1.35$6.75-

Prices updated daily. Last check: Sep 30, 2026

Gemini 3.7 Flash pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
39.6 / 100
Coding
71.5 / 100

Reasoning & Knowledge

  • GPQA Diamond92.1%
  • Humanity's Last Exam39.0%

Coding

  • SciCode59.8%

Agentic & Tool Use

  • Terminal-Bench v2.178.3%
  • τ-bench Banking35.5%

Instruction & Long Context

  • Long-Context Reasoning83.0%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Google
Family
Gemini
Tier
Lightweight
Context Window
1.0M
Modalities
Text, Image, Video, Audio

Capabilities

Tool Calling
Yes
Open Source
No
Aliases
gemini-3.7-flash, Gemini 3.7 Flash, models/gemini-3.7-flash

Strengths & Limitations

Strengths

  • Context window of 1,048,576 tokens supports long documents, large codebases, and extended transcripts in a single request
  • Accepts image, video, and audio inputs in addition to text, so a single model can handle mixed-media pipelines
  • Video and audio input avoid the need for separate transcription or frame-extraction services in many workflows
  • Tool calling is supported, enabling function-calling loops and agent integrations
  • Flash tier is positioned for high-volume, latency-sensitive workloads rather than only occasional complex queries
  • Part of the Gemini family, so prompts and tooling can often be moved between Flash and Pro tiers with limited rework

Limitations

  • As a lightweight-tier model, it is not positioned for the hardest reasoning, math, or long-horizon agentic tasks that Pro-tier Gemini models target
  • We do not currently track published benchmark scores for this model, so capability comparisons rely on your own testing
  • Verified throughput and time-to-first-token measurements are not available in our data
  • Very long prompts approaching the 1M-token limit increase latency and token spend regardless of tier
  • Provider availability for Gemini models is narrower than for open-weight alternatives, limiting deployment options

Key Features

•1,048,576-token (approximately 1M) context window
•Multimodal input: text, images, video, and audio
•Tool calling / function calling support
•Lightweight Flash tier positioned for throughput-oriented workloads
•Part of Google's Gemini model family
•API aliases including gemini-3.7-flash and models/gemini-3.7-flash

About Gemini 3.7 Flash

Gemini 3.7 Flash is a model from Google in the Gemini family, positioned in the Flash line — the lightweight tier Google uses for models aimed at higher-throughput, lower-latency workloads rather than the heaviest reasoning tasks handled by Pro-tier siblings. Like other Flash releases, it is intended for workloads where request volume and response speed matter as much as raw capability. The model carries a context window of 1,048,576 tokens (roughly one million), which allows long documents, large code repositories, extended transcripts, or lengthy chat histories to be placed directly in the prompt rather than retrieved in fragments. It accepts text, image, video, and audio inputs, making it a multimodal model that can process media files alongside written instructions. Tool calling is supported, so it can be wired into function-calling loops, agent frameworks, and structured-output pipelines. In practice, Flash-tier Gemini models are used where a workload runs at scale: document and media ingestion, summarization, extraction, classification, chat assistants, and agent steps that call external tools. Compared with Pro-tier models in the same family, the lightweight tier trades some depth on hard reasoning problems for throughput and cost efficiency. We do not currently track published benchmark scores or verified throughput figures for Gemini 3.7 Flash, so readers evaluating it against peers should benchmark on their own prompts. Availability and pricing vary by provider — see the pricing table on this page.

Common Use Cases

Gemini 3.7 Flash suits workloads that combine large inputs with high request volume: summarizing long reports or contracts, extracting structured fields from documents, transcribing and analyzing audio, describing or indexing video content, classification and routing at scale, and customer-facing chat where response latency matters. The million-token context makes it a candidate for whole-repository code question answering or analyzing long meeting recordings without chunking, while tool calling lets it serve as the execution layer in an agent that queries APIs or databases. For tasks that require deep multi-step reasoning, difficult mathematics, or extended autonomous planning, a Pro-tier Gemini model or another reasoning-focused model is generally the better fit, with Flash reserved for the high-frequency steps around it.

Frequently Asked Questions

How much does Gemini 3.7 Flash cost?

Pricing depends on which provider you use and the pricing model — input versus output tokens, batch versus real-time, and any long-context surcharges. Because rates change frequently, check the pricing table on this page for current per-provider figures rather than relying on a fixed number.

What is Gemini 3.7 Flash best used for?

It is aimed at high-volume, latency-sensitive tasks that still need large inputs: document and transcript summarization, structured data extraction, classification and routing, media understanding across image, video, and audio, and tool-calling agent steps. Save the hardest reasoning tasks for a Pro-tier model.

How does it differ from Gemini Pro-tier models?

Flash is Google's lightweight tier, tuned toward throughput and cost efficiency, while Pro-tier Gemini models target deeper reasoning and more complex agentic work. Both share the Gemini family's long-context and multimodal design, so many prompts port between them, but you should expect Pro tiers to handle harder problems and Flash to handle more requests per unit of spend.

Can Gemini 3.7 Flash process video and audio files?

Yes — our metadata confirms text, image, video, and audio inputs. That allows a single API call to reason over a video clip or audio recording alongside written instructions, which can remove separate transcription or frame-sampling steps from a pipeline.

Does it support function calling for agents?

Tool calling is supported, so the model can be connected to external functions, APIs, and databases within agent frameworks. Specific parameters such as parallel calls or strict schema enforcement depend on the provider's API implementation, so check the provider documentation you plan to use.

How large is the context window in practice?

The context window is 1,048,576 tokens, roughly one million. That is enough for very long documents, multi-hour transcripts, or large portions of a codebase in one prompt, though filling the window increases both latency and token consumption, so most production deployments still trim inputs to what is needed.