Ling-3.0-flash-VL is a chat model from InclusionAI in the Ling 3.0 series, positioned as a speed-oriented "flash" variant with a vision-language (VL) designation.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.060 | $0.180 | $0.012 |
Prices updated daily. Last check: Sep 11, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Ling-3.0-flash-VL fits workloads where generation speed and per-request economy matter more than deep chain-of-thought reasoning: customer-facing chat assistants with streamed output, summarization and rewriting of large document sets, classification and tagging over high request volumes, and content extraction pipelines. Where the provider exposes the vision path implied by the VL designation, it is a candidate for image or screenshot captioning, visual question answering, and document-image understanding at scale. Its ~143 tokens/second output rate supports jobs that generate long responses, while the roughly 1.2-second time to first token means teams building latency-critical voice or autocomplete experiences should benchmark it against alternatives first. For tasks demanding multi-step agentic planning or competition-level math, a higher reasoning tier — within the Ling family or elsewhere — is the more conventional choice.
Pricing depends on which provider hosts the model and on the pricing type — input versus output tokens, and in some cases separate charges for image inputs or batch discounts. Because rates change frequently and differ between endpoints, check the live pricing table on this page for current figures rather than relying on any fixed number.
It suits high-volume, speed-sensitive chat and text-processing work: streamed assistant replies, summarization, rewriting, classification, and extraction pipelines. The VL designation points to vision-language use such as image and document-image understanding where the hosting provider enables that input.
Third-party measurements from Artificial Analysis report roughly 143 output tokens per second with a time to first token of about 1,213 ms. Generation is quick once started, while the initial response delay is above one second, so the model reads as fast for long outputs and less snappy for very short interactive turns. Actual numbers vary by provider, region, and load.
"Flash" indicates the speed- and cost-oriented tier within InclusionAI's Ling 3.0 series rather than the series' heaviest reasoning tier. "VL" marks it as the vision-language variant of that tier, distinguishing it from text-only siblings in the same generation.
We do not have a confirmed context window length for this model in our database. If your workload depends on long inputs, check the model card published by InclusionAI or the documentation of the specific provider you plan to use.