DeepSeek V4 Flash Vision is a chat model from DeepSeek, positioned as the vision-capable variant in the V4 Flash line, with measured output of roughly 117 tokens per second.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.220 | $0.660 | - | |
| $0.220 | $0.660 | $0.0070 | |
| $0.440 | $1.32 | $0.140 | |
| $0.440 | $1.32 | $0.028 |
Prices updated daily. Last check: Sep 5, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
DeepSeek V4 Flash Vision fits workloads that mix images with text and care about response latency: chat assistants that accept screenshots, image captioning and alt-text generation, extraction of fields from photographed forms or receipts, UI and chart description, and moderation-style triage of visual content at volume. Its sub-second time to first token and roughly 117 tokens/second output rate make it a reasonable candidate for user-facing interfaces where partial output appears immediately, and for batch pipelines where per-request wall-clock time drives total job duration. For tasks requiring long chains of reasoning, complex multi-step agentic planning, or heavy code generation, consider evaluating a larger DeepSeek V4 variant or a dedicated reasoning model instead, and validate on your own prompts since we do not track quality benchmarks for this entry.
Pricing depends on the provider, whether you are billed for input or output tokens, and the pricing type offered (on-demand, batch, or committed capacity). Image inputs may also be billed differently from text. Rates change frequently, so check the pricing table on this page for the current per-provider figures rather than relying on a fixed number.
It is aimed at latency-sensitive chat and multimodal workloads — assistants that accept images or screenshots, captioning, visual data extraction, and high-volume request paths where about 117 output tokens per second and a roughly 749 ms time to first token matter more than maximum reasoning depth.
The Flash designation marks it as the throughput-oriented tier within the V4 generation, and the Vision designation marks the variant intended for image-plus-text prompts. Larger V4 siblings generally target deeper reasoning and harder tasks, so a common pattern is to route routine multimodal requests here and escalate the difficult ones.
We do not have a confirmed context window for this model in our database, and served limits can differ by host. Check DeepSeek's model documentation and your provider's API reference before designing around a specific token limit.
The Vision variant name indicates image input alongside text. Specific constraints — supported formats, maximum resolution, and how many images can be attached per request — are not tracked here, so confirm them with the provider you plan to use.