Step 3.7 Flash is a large language model from StepFun, positioned in the Step family's "Flash" line, with measured throughput of roughly 127 output tokens per second.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.200 | $1.15 | $0.040 | |
| $0.200 | $1.15 | $0.040 | |
| $0.200 | $1.15 | $0.040 | |
| $0.200 | $1.15 | $0.040 |
Prices updated daily. Last check: Sep 20, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Step 3.7 Flash fits workloads where generation speed and per-token efficiency matter more than maximum reasoning depth: chat assistants that stream responses to users, content drafting and rewriting, document and conversation summarization, data extraction from unstructured text, and high-volume classification or tagging pipelines. Its measured throughput of roughly 127 output tokens per second supports interactive experiences, while the time to first token of about 1.7 seconds suggests it is a better fit for streaming interfaces — where users see partial output quickly — than for ultra-low-latency single-token responses such as autocomplete. Teams evaluating Chinese-language or bilingual applications may want to include it in a benchmark set alongside models from other families, since StepFun develops primarily for that market. For tasks that demand long multi-step reasoning or verified long-context handling, confirm the model's context window and reasoning benchmarks with the provider before selecting it.
Pricing for Step 3.7 Flash varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are typically billed separately, and different hosts set different rates. See the pricing table on this page for the current per-provider figures.
It suits high-volume text workloads that benefit from fast generation: streaming chat assistants, summarization, content drafting, text extraction, and classification pipelines. The Flash tier positioning and measured throughput of about 127 output tokens per second point toward throughput-oriented rather than maximum-reasoning use.
Artificial Analysis measured approximately 126.7 output tokens per second with a time to first token of roughly 1,743 milliseconds. The throughput figure is well suited to streaming responses; the time to first token is worth testing against your own latency requirements.
We do not currently track a confirmed context window for Step 3.7 Flash. Check StepFun's documentation or your serving provider's API reference for the maximum input length, as hosted deployments sometimes cap context below the model's native limit.
The Flash designation indicates a speed-oriented variant within the Step 3.7 generation, so it is expected to generate faster than heavier siblings while trading off some capability. Because we do not have comparable benchmark scores across the family in our dataset, running your own evaluation on representative prompts is the most reliable way to compare them.
Step 3.7 Flash is developed by StepFun, a Chinese AI company behind the Step series of large language models.