Mistral Large 4 Preview
Mistral Large 4 Preview is a preview release of Mistral's Large-series chat model, offered through inference APIs while the generally available version is finalized.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.680 | $2.09 | $0.070 |
Prices updated daily. Last check: Oct 7, 2026
Mistral Large 4 Preview pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- Humanity's Last Exam35.0%
Coding
- SciCode54.2%
Instruction & Long Context
- Long-Context Reasoning81.3%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Mistral
- Modalities
- Text
Capabilities
Strengths & Limitations
Strengths
- Part of Mistral's Large series, the company's general-purpose line above its compact Ministral and specialized Codestral/Pixtral models
- Measured time to first token of roughly 971 ms, low enough for streaming chat UIs where perceived responsiveness matters
- Measured output throughput of approximately 106 tokens per second in Artificial Analysis testing
- Preview availability lets teams begin integration and evaluation work before the stable release lands
- Mistral models are commonly served by multiple inference providers, so cost and latency can be compared side by side in the pricing table
- European-developed model line, which matters to organizations with data-residency or vendor-diversity requirements
Limitations
- Preview release — behavior, API surface, and provider availability may change before a stable version ships
- We do not have a confirmed context window figure for this build, so long-document workloads need verification against the provider's docs
- Modality support (for example image input) is not confirmed in our data and should be checked with the serving provider
- Published benchmark scores for reasoning, coding, or multilingual tasks are not available in our dataset for this preview
- Throughput and first-token latency vary by provider and load; the measured figures may not match a specific endpoint
Key Features
About Mistral Large 4 Preview
Common Use Cases
Mistral Large 4 Preview is aimed at general-purpose assistant workloads — multi-turn chat, drafting and rewriting, summarization, question answering over supplied context, and the orchestration layers of application backends — where a Large-tier model's broader capability is preferred over a compact model. Its sub-second measured time to first token suits interactive, user-facing surfaces where the first words appearing quickly matters more than total completion time, while the ~106 tokens/second output rate is adequate for medium-length responses but is worth benchmarking against alternatives for workloads that generate very long outputs or run large batch jobs. Because it is a preview build, it fits best in evaluation harnesses, prototypes, and internal tooling where teams want early access and can tolerate changes; production systems with strict stability requirements may prefer a generally available Mistral release until the stable Large 4 version ships.
Frequently Asked Questions
How much does Mistral Large 4 Preview cost?
Pricing depends on which inference provider serves the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and differ between providers. Rates also change frequently. Check the pricing table on this page for the current per-provider figures.
What is Mistral Large 4 Preview best used for?
It suits general-purpose chat and assistant workloads: multi-turn conversation, drafting and editing text, summarization, and question answering against supplied context. The measured ~971 ms time to first token makes it reasonable for streaming interfaces, and as a Large-series model it is positioned for broader tasks than Mistral's compact Ministral line.
What does the "Preview" in the name mean?
It signals an early release made available for evaluation ahead of a stable, generally available version. In practice that means model behavior, API details, and which providers host it can change, so it is better suited to prototyping and benchmarking than to systems that require a frozen model version.
How fast is Mistral Large 4 Preview?
Artificial Analysis measurements put it at roughly 106 output tokens per second with a time to first token of about 971 ms. Those are reference figures — actual speed varies by provider, region, prompt length, and current load, so benchmark your own endpoint before committing to latency targets.
What context window does it support?
We do not have a confirmed context window figure for this preview build in our database, and we would rather say so than guess. Check the serving provider's model documentation for the exact token limit, since providers occasionally cap context below a model's native maximum.