Skip to main content
Open SourceDeepSeek

DeepSeek V4 Flash Vision

DeepSeek V4 Flash Vision is a chat model from DeepSeek, positioned as the vision-capable variant in the V4 Flash line, with measured output of roughly 117 tokens per second.

License Open Source
Input from
$0.216 / 1M tokens
across 3 providers

API Pricing

Cheapest on OpenRouter — 41% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.216$0.647$0.0069
$0.440$1.32$0.014
$0.440$1.32$0.014

Prices updated daily. Last check: Oct 11, 2026

Compare API pricing for every DeepSeek model →

DeepSeek V4 Flash Vision pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
34.8 / 100
Coding
65.0 / 100
Output Speed
213 t/s
Latency (TTFT)
933ms

Reasoning & Knowledge

  • GPQA Diamond91.3%
  • Humanity's Last Exam34.5%

Coding

  • SciCode49.7%

Agentic & Tool Use

  • Terminal-Bench v2.174.2%
  • τ-bench Banking41.0%

Instruction & Long Context

  • Long-Context Reasoning81.3%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
DeepSeek
Modalities
Text

Capabilities

Open Source
Yes

Strengths & Limitations

Strengths

  • Measured output throughput of about 117 tokens per second in Artificial Analysis testing, suitable for streaming chat responses
  • Time to first token of roughly 749 ms keeps interactive turn-taking responsive
  • Vision-labeled variant, intended for prompts that combine images with text
  • Positioned in the Flash tier of DeepSeek's V4 generation, which targets latency- and cost-sensitive workloads rather than maximum reasoning depth
  • Part of a broader DeepSeek V4 family, allowing tiered routing where harder requests can be escalated to a heavier sibling
  • Available through third-party inference providers, so pricing and throughput can be compared side by side

Limitations

  • We do not have a confirmed context window for this model, so long-document limits should be verified with the provider
  • Flash-tier models generally trade reasoning depth for speed compared with larger siblings in the same generation
  • Published benchmark coverage in our database is limited to speed and latency, with no quality scores to compare against peers
  • Image input constraints such as maximum resolution, images per request, and supported formats are not tracked here
  • Serving configuration and speed can differ between providers, so measured throughput may not match what you observe

Key Features

•Chat completion interface for multi-turn conversations
•Image input support implied by the Vision variant designation
•Flash tier of the DeepSeek V4 model generation
•Measured output speed of approximately 117 tokens/second (Artificial Analysis)
•Measured time to first token of approximately 749 ms (Artificial Analysis)
•Streaming responses for incremental output delivery
•Multiple hosting options tracked on this page for price and performance comparison

About DeepSeek V4 Flash Vision

DeepSeek V4 Flash Vision is a chat-oriented model released by DeepSeek as part of its V4 generation. The "Flash" designation places it in the throughput-oriented segment of that generation rather than the heavier reasoning tier, and the "Vision" designation indicates the variant intended for prompts that include image input alongside text. Independent measurements collected by Artificial Analysis put the model's output speed at about 117 tokens per second with a time to first token of roughly 749 ms. Together these numbers describe a model that begins responding under a second in typical conditions and streams at a rate suitable for interactive chat and streaming UI patterns. We do not currently track a confirmed context window, image resolution limits, or tool-calling details for this entry, so check DeepSeek's own model documentation and your chosen provider's API reference before building against specific limits. In practice, models in a "flash" or fast tier are used where response latency and per-request cost matter more than maximum reasoning depth: chat assistants, document and screenshot understanding, image captioning and extraction pipelines, and high-volume request paths. Buyers comparing DeepSeek V4 Flash Vision against other multimodal chat models should weigh its measured latency and throughput against the depth of larger DeepSeek V4 variants and against multimodal offerings from other creators, since availability and served configuration vary by provider.

Common Use Cases

DeepSeek V4 Flash Vision fits workloads that mix images with text and care about response latency: chat assistants that accept screenshots, image captioning and alt-text generation, extraction of fields from photographed forms or receipts, UI and chart description, and moderation-style triage of visual content at volume. Its sub-second time to first token and roughly 117 tokens/second output rate make it a reasonable candidate for user-facing interfaces where partial output appears immediately, and for batch pipelines where per-request wall-clock time drives total job duration. For tasks requiring long chains of reasoning, complex multi-step agentic planning, or heavy code generation, consider evaluating a larger DeepSeek V4 variant or a dedicated reasoning model instead, and validate on your own prompts since we do not track quality benchmarks for this entry.

Frequently Asked Questions

How much does DeepSeek V4 Flash Vision cost?

Pricing depends on the provider, whether you are billed for input or output tokens, and the pricing type offered (on-demand, batch, or committed capacity). Image inputs may also be billed differently from text. Rates change frequently, so check the pricing table on this page for the current per-provider figures rather than relying on a fixed number.

What is DeepSeek V4 Flash Vision best used for?

It is aimed at latency-sensitive chat and multimodal workloads — assistants that accept images or screenshots, captioning, visual data extraction, and high-volume request paths where about 117 output tokens per second and a roughly 749 ms time to first token matter more than maximum reasoning depth.

How does it differ from other DeepSeek V4 models?

The Flash designation marks it as the throughput-oriented tier within the V4 generation, and the Vision designation marks the variant intended for image-plus-text prompts. Larger V4 siblings generally target deeper reasoning and harder tasks, so a common pattern is to route routine multimodal requests here and escalate the difficult ones.

What context window does DeepSeek V4 Flash Vision support?

We do not have a confirmed context window for this model in our database, and served limits can differ by host. Check DeepSeek's model documentation and your provider's API reference before designing around a specific token limit.

Can it accept image input?

The Vision variant name indicates image input alongside text. Specific constraints — supported formats, maximum resolution, and how many images can be attached per request — are not tracked here, so confirm them with the provider you plan to use.