Skip to main content
Mistral

Mistral Large 4 Preview

Mistral Large 4 Preview is a preview release of Mistral's Large-series chat model, offered through inference APIs while the generally available version is finalized.

Input from
$0.680 / 1M tokens
across 1 provider

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.680$2.09$0.070

Prices updated daily. Last check: Oct 7, 2026

Mistral Large 4 Preview pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
38.4 / 100
Output Speed
106 t/s
Latency (TTFT)
971ms

Reasoning & Knowledge

  • Humanity's Last Exam35.0%

Coding

  • SciCode54.2%

Instruction & Long Context

  • Long-Context Reasoning81.3%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Mistral
Modalities
Text

Capabilities

Strengths & Limitations

Strengths

  • Part of Mistral's Large series, the company's general-purpose line above its compact Ministral and specialized Codestral/Pixtral models
  • Measured time to first token of roughly 971 ms, low enough for streaming chat UIs where perceived responsiveness matters
  • Measured output throughput of approximately 106 tokens per second in Artificial Analysis testing
  • Preview availability lets teams begin integration and evaluation work before the stable release lands
  • Mistral models are commonly served by multiple inference providers, so cost and latency can be compared side by side in the pricing table
  • European-developed model line, which matters to organizations with data-residency or vendor-diversity requirements

Limitations

  • Preview release — behavior, API surface, and provider availability may change before a stable version ships
  • We do not have a confirmed context window figure for this build, so long-document workloads need verification against the provider's docs
  • Modality support (for example image input) is not confirmed in our data and should be checked with the serving provider
  • Published benchmark scores for reasoning, coding, or multilingual tasks are not available in our dataset for this preview
  • Throughput and first-token latency vary by provider and load; the measured figures may not match a specific endpoint

Key Features

•Chat/instruction-following model in Mistral's Large series
•Measured ~106 output tokens per second (Artificial Analysis)
•Measured ~971 ms time to first token (Artificial Analysis)
•Preview release channel for pre-GA evaluation
•Served through inference APIs with per-provider pricing comparison
•Streaming response support typical of Mistral chat endpoints

About Mistral Large 4 Preview

Mistral Large 4 Preview is a chat model from Mistral, the French AI lab known for the Mistral, Mixtral, Ministral, Codestral, and Pixtral model lines. It sits in the Large series, which Mistral uses for its higher-capability general-purpose assistants rather than its compact or task-specialized models. The "Preview" label indicates an early release intended for evaluation and integration testing ahead of a stable version, so behavior, availability, and provider coverage can shift during the preview window. In third-party measurements from Artificial Analysis, Mistral Large 4 Preview has been observed generating roughly 106 output tokens per second with a time to first token of about 971 ms. That combination puts it in the mid-range for interactive use: the sub-second first-token latency is workable for streaming chat interfaces, while the sustained output rate affects how quickly long responses complete. Both figures are provider- and load-dependent, so the numbers in our benchmark data should be treated as a reference point rather than a guarantee for any particular endpoint. Beyond those measured throughput and latency figures, we do not track confirmed context window, modality, or benchmark-score details for this preview build, and we have deliberately left those out rather than estimate them. Readers evaluating the model for production should check the serving provider's own model card for the context limit, supported input types, and tool-calling support, and compare the per-provider entries in the pricing table on this page for current cost and availability.

Common Use Cases

Mistral Large 4 Preview is aimed at general-purpose assistant workloads — multi-turn chat, drafting and rewriting, summarization, question answering over supplied context, and the orchestration layers of application backends — where a Large-tier model's broader capability is preferred over a compact model. Its sub-second measured time to first token suits interactive, user-facing surfaces where the first words appearing quickly matters more than total completion time, while the ~106 tokens/second output rate is adequate for medium-length responses but is worth benchmarking against alternatives for workloads that generate very long outputs or run large batch jobs. Because it is a preview build, it fits best in evaluation harnesses, prototypes, and internal tooling where teams want early access and can tolerate changes; production systems with strict stability requirements may prefer a generally available Mistral release until the stable Large 4 version ships.

Frequently Asked Questions

How much does Mistral Large 4 Preview cost?

Pricing depends on which inference provider serves the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and differ between providers. Rates also change frequently. Check the pricing table on this page for the current per-provider figures.

What is Mistral Large 4 Preview best used for?

It suits general-purpose chat and assistant workloads: multi-turn conversation, drafting and editing text, summarization, and question answering against supplied context. The measured ~971 ms time to first token makes it reasonable for streaming interfaces, and as a Large-series model it is positioned for broader tasks than Mistral's compact Ministral line.

What does the "Preview" in the name mean?

It signals an early release made available for evaluation ahead of a stable, generally available version. In practice that means model behavior, API details, and which providers host it can change, so it is better suited to prototyping and benchmarking than to systems that require a frozen model version.

How fast is Mistral Large 4 Preview?

Artificial Analysis measurements put it at roughly 106 output tokens per second with a time to first token of about 971 ms. Those are reference figures — actual speed varies by provider, region, prompt length, and current load, so benchmark your own endpoint before committing to latency targets.

What context window does it support?

We do not have a confirmed context window figure for this preview build in our database, and we would rather say so than guess. Check the serving provider's model documentation for the exact token limit, since providers occasionally cap context below a model's native maximum.