Mistral Large 3 is a large multimodal model from Mistral in the Mistral Large family, accepting text and image input with a 256K token context window.
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.250 | $0.750 | |
| $0.500 | $1.50 | |
| $0.500 | $1.50 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Mistral Large 3 fits workloads that need a large-tier model rather than a cheap classifier: long-document summarization and question answering, contract and report analysis where the 256K context lets an entire file stay in the prompt, multilingual drafting and translation, and code generation or review across many files. Image input extends it to mixed workloads such as reading scanned pages, screenshots, diagrams, or UI mockups together with instructions. Its measured throughput of around 59 tokens per second suits assistant and back-office automation more than ultra-low-latency streaming interfaces; for very high request volumes or simple extraction and routing tasks, a smaller Mistral model will usually be more cost-effective, with Mistral Large 3 reserved for the steps that need more capability.
Pricing depends on the provider and the pricing model — per-input-token and per-output-token rates differ between hosts, and some offer batch or committed-capacity rates. Check the pricing table on this page for current per-provider figures, and remember that long 256K-context prompts consume many input tokens per request.
It suits long-context document analysis, multilingual writing and translation, code generation and review, and multi-step agent workflows where text and images arrive together. The 256K window means large files or long chat histories can stay in a single prompt instead of being chunked.
Yes. Our metadata lists both text and image input modalities, so images can be supplied alongside text in a prompt. Exact limits on image count, resolution, and formats depend on the hosting provider's API.
Artificial Analysis measurements record about 59 output tokens per second with a time to first token near 652 ms. Actual figures vary by provider, region, prompt length, and concurrency, so treat these as a general reference.
It is the dated version identifier for this release of Mistral Large 3 (year 25, month 12). Using the dated ID rather than a rolling alias pins your application to a specific model version.
Mistral Large 3 sits in the Large tier and is aimed at harder reasoning, coding, and long-context tasks. For high-volume classification, routing, or short extraction jobs, a smaller model in Mistral's lineup will typically be cheaper and faster; compare both in the pricing table if you plan to route requests between tiers.