Ling-3.0-flash-Fin is a chat model from InclusionAI in the Ling 3.0 series, positioned as a "flash" speed-oriented variant with a domain-specific "Fin" designation.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.060 | $0.180 | $0.012 | |
| $0.060 | $0.180 | $0.012 |
Prices updated daily. Last check: Sep 17, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Ling-3.0-flash-Fin fits workloads with a narrow, predictable prompt distribution in the financial domain — summarizing filings and reports, answering questions over financial documents, drafting analyst notes, extracting structured fields from financial text, and powering internal assistants for finance teams. Its measured throughput of roughly 163 output tokens per second makes it practical for batch document processing where many generations must complete within a fixed window, and the sub-2-second time to first token is adequate for chat interfaces where a short pause before streaming begins is acceptable. For general-purpose assistants, broad coding work, or tasks far outside the financial domain, a general Ling 3.0 variant or another general-purpose chat model is the more appropriate starting point. Because our dataset carries only serving performance for this model, run a head-to-head evaluation on your own financial-text test set before standardizing on it.
Pricing depends on which inference provider serves the model and on the pricing type — input tokens, output tokens, and any cached-input rates are billed separately, and providers change rates over time. See the pricing table on this page for current per-provider rates rather than relying on any figure quoted elsewhere.
It is best suited to financial-domain text work — document question answering, report and filing summarization, field extraction from financial text, and finance-team assistants — where its "Fin" specialization is relevant and its ~163 tokens/second output throughput helps with batch or interactive volume.
The "flash" label marks it as a speed-oriented configuration within the Ling 3.0 series, and the "Fin" suffix marks it as domain-targeted rather than general-purpose. Other Ling 3.0 variants are the better comparison point for broad, general-purpose chat workloads.
Artificial Analysis measured approximately 163 output tokens per second with a time to first token of about 1.56 seconds. Actual figures vary by provider, region, prompt length, and load, so test against your own traffic pattern.
We do not currently track a confirmed context window for this model. Check the documentation of the provider you plan to use, since served context limits can also be lower than the model's native maximum.