DeepSeek V4 Flash Vision
DeepSeek V4 Flash Vision is a chat model from DeepSeek, positioned as the vision-capable variant in the V4 Flash line, with measured output of roughly 117 tokens per second.
API Pricing
Cheapest on OpenRouter — 41% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.216 | $0.647 | $0.0069 | |
| $0.440 | $1.32 | $0.014 | |
| $0.440 | $1.32 | $0.014 |
Prices updated daily. Last check: Oct 11, 2026
Compare API pricing for every DeepSeek model →DeepSeek V4 Flash Vision pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond91.3%
- Humanity's Last Exam34.5%
Coding
- SciCode49.7%
Agentic & Tool Use
- Terminal-Bench v2.174.2%
- τ-bench Banking41.0%
Instruction & Long Context
- Long-Context Reasoning81.3%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- DeepSeek
- Modalities
- Text
Capabilities
- Open Source
- Yes
Strengths & Limitations
Strengths
- Measured output throughput of about 117 tokens per second in Artificial Analysis testing, suitable for streaming chat responses
- Time to first token of roughly 749 ms keeps interactive turn-taking responsive
- Vision-labeled variant, intended for prompts that combine images with text
- Positioned in the Flash tier of DeepSeek's V4 generation, which targets latency- and cost-sensitive workloads rather than maximum reasoning depth
- Part of a broader DeepSeek V4 family, allowing tiered routing where harder requests can be escalated to a heavier sibling
- Available through third-party inference providers, so pricing and throughput can be compared side by side
Limitations
- We do not have a confirmed context window for this model, so long-document limits should be verified with the provider
- Flash-tier models generally trade reasoning depth for speed compared with larger siblings in the same generation
- Published benchmark coverage in our database is limited to speed and latency, with no quality scores to compare against peers
- Image input constraints such as maximum resolution, images per request, and supported formats are not tracked here
- Serving configuration and speed can differ between providers, so measured throughput may not match what you observe
Key Features
About DeepSeek V4 Flash Vision
Common Use Cases
DeepSeek V4 Flash Vision fits workloads that mix images with text and care about response latency: chat assistants that accept screenshots, image captioning and alt-text generation, extraction of fields from photographed forms or receipts, UI and chart description, and moderation-style triage of visual content at volume. Its sub-second time to first token and roughly 117 tokens/second output rate make it a reasonable candidate for user-facing interfaces where partial output appears immediately, and for batch pipelines where per-request wall-clock time drives total job duration. For tasks requiring long chains of reasoning, complex multi-step agentic planning, or heavy code generation, consider evaluating a larger DeepSeek V4 variant or a dedicated reasoning model instead, and validate on your own prompts since we do not track quality benchmarks for this entry.
Frequently Asked Questions
How much does DeepSeek V4 Flash Vision cost?
Pricing depends on the provider, whether you are billed for input or output tokens, and the pricing type offered (on-demand, batch, or committed capacity). Image inputs may also be billed differently from text. Rates change frequently, so check the pricing table on this page for the current per-provider figures rather than relying on a fixed number.
What is DeepSeek V4 Flash Vision best used for?
It is aimed at latency-sensitive chat and multimodal workloads — assistants that accept images or screenshots, captioning, visual data extraction, and high-volume request paths where about 117 output tokens per second and a roughly 749 ms time to first token matter more than maximum reasoning depth.
How does it differ from other DeepSeek V4 models?
The Flash designation marks it as the throughput-oriented tier within the V4 generation, and the Vision designation marks the variant intended for image-plus-text prompts. Larger V4 siblings generally target deeper reasoning and harder tasks, so a common pattern is to route routine multimodal requests here and escalate the difficult ones.
What context window does DeepSeek V4 Flash Vision support?
We do not have a confirmed context window for this model in our database, and served limits can differ by host. Check DeepSeek's model documentation and your provider's API reference before designing around a specific token limit.
Can it accept image input?
The Vision variant name indicates image input alongside text. Specific constraints — supported formats, maximum resolution, and how many images can be attached per request — are not tracked here, so confirm them with the provider you plan to use.