Gemini 3.5 Flash is a model in Google's Gemini Flash line, the family's speed-and-efficiency tier, measured at roughly 260 output tokens per second in third-party benchmarking.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.600 | $3.60 | $0.060 | |
| $0.750 | $4.50 | $0.075 | |
| $0.750 | $4.50 | $0.075 | |
| $1.50 | $9.00 | - | |
| $1.50 | $9.00 | $0.150 | |
| $1.50 | $9.00 | $0.150 | |
| $1.50 | $9.00 | $0.150 |
Prices updated daily. Last check: Sep 21, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Gemini 3.5 Flash's measured profile — high output token throughput with a long time to first token — points toward batch and long-output workloads rather than snappy conversational turns. Good candidates include document summarization, bulk content generation, data extraction and transformation across large record sets, and background agent steps where a few seconds of initial delay is acceptable but total generation time matters. As a Flash-tier model it is the kind of option teams typically reach for when running high request volumes and wanting to reserve larger Gemini models for the subset of requests that need deeper reasoning. For chat interfaces where perceived responsiveness on short replies is the priority, the measured first-token latency is a consideration worth testing against your own prompts.
Pricing varies by provider and by pricing type — input versus output tokens, and whether a provider offers batch or cached-input rates. Rates also change frequently. Check the pricing table on this page for the current figures from each provider we track.
Its measured throughput of roughly 260 output tokens per second suits long-output and high-volume work: summarization, bulk extraction and classification, content generation, and background agent steps. The measured time to first token of about 11.6 seconds makes it less suited to interactive chat where a fast first word matters.
It belongs to Google's Flash line, which within the Gemini family is the tier oriented toward speed and efficiency rather than maximum capability. Larger Gemini tiers are generally the choice for the hardest reasoning and agentic tasks, while Flash models target volume workloads. Compare the specific models listed on this site for the measurements we have for each.
We do not currently have a confirmed context window figure for this model in our database. Google's official model documentation is the authoritative source for context limits and any input modality support.
A long time to first token alongside a fast decode rate is a pattern commonly associated with models that deliberate internally before producing visible output. The practical effect is that short responses feel slow while long responses complete comparatively quickly. Both figures come from Artificial Analysis measurements and will vary with prompt length, provider, and load.