Skip to main content
Mistral

Mistral Large 3

Mistral Large 3 is a large multimodal model from Mistral in the Mistral Large family, accepting text and image input with a 256K token context window.

Context 256K
Modalities text, image
Input from
$0.250 / 1M tokens
across 2 providers

API Pricing

Cheapest on Amazon AWS — 40% below avg
ProviderInput / 1MOutput / 1M
$0.250$0.750
$0.500$1.50
$0.500$1.50

Prices updated daily. Last check: Sep 24, 2026

Mistral Large 3 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
9.3 / 100
Coding
20.1 / 100
Math
38.0 / 100
Output Speed
73.2 t/s
Latency (TTFT)
751ms

Reasoning & Knowledge

  • MMLU-Pro80.7%
  • GPQA Diamond68.0%
  • Humanity's Last Exam4.2%

Coding

  • LiveCodeBench46.5%
  • SciCode36.6%

Math

  • AIME 202538.0%

Agentic & Tool Use

  • Terminal-Bench Hard15.9%
  • Terminal-Bench v2.112.0%
  • τ²-bench24.6%
  • τ-bench Banking5.8%

Instruction & Long Context

  • IFBench36.2%
  • Long-Context Reasoning36.0%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Mistral
Family
Mistral Large
Context Window
256K
Modalities
Text, Image

Capabilities

Tool Calling
No
Open Source
No
Aliases
mistralai/Mistral-Large-3, mistral-large-3, mistral-large-2512, Mistral Large 3

Strengths & Limitations

Strengths

  • 256,000 token context window, enough for long documents, large codebases, or extended multi-turn sessions in one request
  • Accepts image input alongside text, so visual documents, screenshots, and charts can be handled in the same call as text
  • Measured output speed of about 59 tokens per second and time to first token near 652 ms in Artificial Analysis testing
  • Available under multiple provider aliases (mistral-large-3, mistral-large-2512), making it easier to find across hosting options
  • Dated version identifier (2512) allows pinning to a specific release rather than a moving alias
  • Positioned in Mistral's Large tier, so it targets harder reasoning and generation tasks than the company's smaller models

Limitations

  • Roughly 59 output tokens per second is moderate — latency-critical, high-volume workloads may be better served by a smaller model
  • Sub-second time to first token is adequate but not the fastest available for interactive UIs
  • Large-tier pricing is generally higher per token than compact models, so long 256K-context prompts can become expensive
  • We do not track detailed feature support (for example tool calling specifics or output modalities) for this model, so verify against your provider's documentation
  • Benchmark scores for reasoning, coding, and math are not present in our data set, making direct quality comparisons with peers difficult from this page alone

Key Features

•256K token context window
•Text and image input (multimodal prompting)
•Mistral Large family, third generation
•Dated release identifier: mistral-large-2512
•Multiple provider aliases including mistralai/Mistral-Large-3
•Measured throughput ~59 output tokens/sec (Artificial Analysis)
•Measured time to first token ~652 ms (Artificial Analysis)
•Available from multiple inference providers for price comparison

About Mistral Large 3

Mistral Large 3 is a model from the French AI developer Mistral, and the third generation in its Mistral Large line — the company's large, general-purpose tier. It is versioned as mistral-large-2512 and appears across providers under aliases including mistral-large-3 and mistralai/Mistral-Large-3, so the same model may be listed under several identifiers in the pricing table on this page. Mistral Large 3 accepts both text and image input and supports a context window of 256,000 tokens, which allows long documents, extended chat histories, large code contexts, or batches of images and accompanying text to be processed in a single request. Independent measurements from Artificial Analysis put its output speed at roughly 59 tokens per second with a time to first token of about 652 ms; both figures vary by hosting provider, region, and load, so treat them as a reference point rather than a guarantee for any specific endpoint. In practice, models at this tier are used for work where output quality matters more than raw throughput: document analysis, multilingual assistants, code generation and review, and agent-style pipelines that chain several steps. Compared with smaller Mistral models, Mistral Large 3 sits at the larger end of the lineup, and its 256K context and image input make it a candidate for mixed text-and-visual workloads that would otherwise require splitting inputs across requests.

Common Use Cases

Mistral Large 3 fits workloads that need a large-tier model rather than a cheap classifier: long-document summarization and question answering, contract and report analysis where the 256K context lets an entire file stay in the prompt, multilingual drafting and translation, and code generation or review across many files. Image input extends it to mixed workloads such as reading scanned pages, screenshots, diagrams, or UI mockups together with instructions. Its measured throughput of around 59 tokens per second suits assistant and back-office automation more than ultra-low-latency streaming interfaces; for very high request volumes or simple extraction and routing tasks, a smaller Mistral model will usually be more cost-effective, with Mistral Large 3 reserved for the steps that need more capability.

Frequently Asked Questions

How much does Mistral Large 3 cost?

Pricing depends on the provider and the pricing model — per-input-token and per-output-token rates differ between hosts, and some offer batch or committed-capacity rates. Check the pricing table on this page for current per-provider figures, and remember that long 256K-context prompts consume many input tokens per request.

What is Mistral Large 3 best used for?

It suits long-context document analysis, multilingual writing and translation, code generation and review, and multi-step agent workflows where text and images arrive together. The 256K window means large files or long chat histories can stay in a single prompt instead of being chunked.

Can Mistral Large 3 process images?

Yes. Our metadata lists both text and image input modalities, so images can be supplied alongside text in a prompt. Exact limits on image count, resolution, and formats depend on the hosting provider's API.

How fast is Mistral Large 3?

Artificial Analysis measurements record about 59 output tokens per second with a time to first token near 652 ms. Actual figures vary by provider, region, prompt length, and concurrency, so treat these as a general reference.

What does mistral-large-2512 mean?

It is the dated version identifier for this release of Mistral Large 3 (year 25, month 12). Using the dated ID rather than a rolling alias pins your application to a specific model version.

Should I choose Mistral Large 3 or a smaller Mistral model?

Mistral Large 3 sits in the Large tier and is aimed at harder reasoning, coding, and long-context tasks. For high-volume classification, routing, or short extraction jobs, a smaller model in Mistral's lineup will typically be cheaper and faster; compare both in the pricing table if you plan to route requests between tiers.