DeepSeek V4.1 Flash is a chat model from DeepSeek positioned as a speed-oriented variant in the company's V4.1 line, measured at roughly 267 output tokens per second.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.150 | $0.600 | $0.0030 | |
| $0.200 | $0.600 | $0.0060 | |
| $0.254 | $0.972 | $0.127 | |
| $0.300 | $1.20 | $0.0060 | |
| $0.300 | $1.20 | $0.0060 |
Prices updated daily. Last check: Sep 11, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
DeepSeek V4.1 Flash fits workloads defined by request volume rather than problem difficulty: customer-facing chat assistants where streaming responsiveness shapes the user experience, document and transcript summarization pipelines, content rewriting and translation batches, retrieval-augmented question answering over indexed corpora, and routing or classification steps inside a larger agent system. Its measured throughput of roughly 267 tokens per second means long generations complete in a few seconds, which matters for drafting and summarization jobs, while sub-second time to first token keeps interactive sessions feeling immediate. For work requiring extended chain-of-thought reasoning, competition-level mathematics, or long autonomous coding sessions, a larger reasoning-oriented model — including heavier members of the DeepSeek lineup — is usually the better match, with Flash reserved for the high-frequency, lower-complexity calls in the same application.
Pricing depends on which provider hosts the model and on the pricing type — input tokens, output tokens, cached input, and any batch or committed-throughput discounts are billed separately. Rates also change as providers compete. Check the pricing table on this page for the current figures across the providers we track.
High-volume text workloads where speed and per-request cost dominate: chat assistants, summarization, translation, classification, extraction, and retrieval-augmented answering. Its measured ~267 tokens/second output rate and ~896 ms time to first token make it a reasonable fit for interactive and batch pipelines alike.
The Flash designation indicates the speed- and cost-oriented tier of the V4.1 generation. Non-Flash siblings are generally positioned for harder reasoning and longer-horizon tasks, while Flash targets fast responses across many requests. If your workload mixes both, a common pattern is routing simple calls to Flash and escalating difficult ones to a larger model.
We do not have a confirmed context window recorded for this model. Providers sometimes serve the same DeepSeek model with different maximum context settings, so consult the model card of the specific provider you plan to use.
Independent measurement puts time to first token at roughly 896 ms and sustained output at about 267 tokens per second. That is fast enough for streaming chat interfaces, where the first token arrives in under a second and text renders faster than most people read, though it is slower to start than latency-specialized inference stacks.