Skip to main content
Open SourceMistral

Ministral 3 3B

Ministral 3 3B is a small chat model from Mistral, part of the Ministral line of compact models aimed at high-throughput, latency-sensitive text workloads.

License Open Source
Input from
$0.050 / 1M tokens
across 1 provider

API Pricing

Cheapest on Amazon AWS — 33% below avg
ProviderInput / 1MOutput / 1M
$0.050$0.050
$0.100$0.100

Prices updated daily. Last check: Oct 10, 2026

Compare API pricing for every Mistral model →

Ministral 3 3B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
4.8 / 100
Coding
4.8 / 100
Math
22.0 / 100
Output Speed
226 t/s
Latency (TTFT)
541ms

Reasoning & Knowledge

  • MMLU-Pro52.4%
  • GPQA Diamond35.8%
  • Humanity's Last Exam5.4%

Coding

  • LiveCodeBench24.7%
  • SciCode15.3%

Math

  • AIME 202522.0%

Agentic & Tool Use

  • Terminal-Bench Hard0.0%
  • Terminal-Bench v2.10.0%
  • τ²-bench24.9%
  • τ-bench Banking4.7%

Instruction & Long Context

  • IFBench26.8%
  • Long-Context Reasoning17.0%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Mistral
Modalities
Text

Capabilities

Open Source
Yes

Strengths & Limitations

Strengths

  • Compact ~3B parameter scale, which generally translates to lower cost per token than larger Mistral models
  • Measured output throughput of about 201 tokens per second in Artificial Analysis testing
  • Time to first token measured at roughly 456 ms, suitable for interactive chat turns
  • Part of Mistral's Ministral small-model line, so it shares prompt conventions with other Mistral models for easy swapping
  • Small size makes it a practical default tier in cascaded routing setups where a larger model handles escalations
  • Throughput profile fits high-volume batch text jobs such as tagging, extraction, and summarization

Limitations

  • A ~3B parameter model will generally trail larger Mistral models on multi-step reasoning, long-form coding, and nuanced instruction following
  • We do not currently track a confirmed context window for this model — verify with your provider before long-document use
  • Multimodal input support is not tracked in our data; do not assume image or audio handling
  • Reasoning-heavy benchmark scores are not available in our dataset, making direct quality comparisons with peers harder
  • Provider availability for small Mistral models is typically narrower than for the larger flagship models

Key Features

•Chat/instruct-style text generation
•Compact ~3B parameter class model
•Measured ~201 output tokens/second throughput (Artificial Analysis)
•Measured ~456 ms time to first token (Artificial Analysis)
•Member of Mistral's Ministral compact model family
•Suited to cascaded/routing architectures as the low-cost tier
•Multi-provider price comparison available in the table on this page

About Ministral 3 3B

Ministral 3 3B is a chat model from Mistral, positioned in the company's Ministral family — the naming Mistral uses for its smaller, compact-scale models rather than its larger Mistral or Magistral lines. The "3B" in the name refers to the roughly 3-billion-parameter scale, placing it at the lightweight end of the catalog where cost per token and response latency usually matter more than maximum reasoning depth. On measured serving performance, third-party benchmarking from Artificial Analysis records roughly 201 output tokens per second with a time to first token of about 456 ms. Those figures are consistent with what a compact model is generally used for: interactive assistants where the first token needs to arrive quickly, and batch text processing where total tokens generated per dollar and per second dominate the decision. We do not currently track a confirmed context window, modality list, or tool-calling support for this entry, so check your provider's model card before depending on those features. In practice, models at this scale are selected for the high-volume half of a two-model setup: Ministral 3 3B handles routine classification, extraction, rewriting, and short conversational turns, while a larger Mistral model is reserved for the requests that fail a confidence check. Because several providers may serve the same weights, throughput and latency can differ from the benchmark figures above depending on which endpoint you use — compare the providers listed in the pricing table on this page.

Common Use Cases

Ministral 3 3B fits workloads where request volume and response latency drive the architecture more than peak reasoning quality. Typical fits include intent and sentiment classification, structured field extraction from short documents, message drafting and rewriting, autocomplete and suggestion features, chat assistants with narrow scope, and content routing or triage before a larger model is invoked. Its measured sub-half-second time to first token makes it a reasonable choice for streaming interfaces where perceived responsiveness matters, and its ~201 tokens/second output rate supports batch pipelines that generate large volumes of short completions. For tasks involving long multi-step reasoning chains, complex codebase work, or long-context document analysis, a larger model in the Mistral lineup is usually the better match, with Ministral 3 3B handling the routine share of traffic.

Frequently Asked Questions

How much does Ministral 3 3B cost to use?

Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and change frequently. See the pricing table on this page for current per-provider rates rather than relying on a fixed figure.

What is Ministral 3 3B best used for?

It is best suited to high-volume, latency-sensitive text tasks: classification, extraction, rewriting, short chat turns, and triage steps in a larger pipeline. Its compact ~3B scale and measured ~456 ms time to first token favor throughput and responsiveness over deep multi-step reasoning.

How fast is Ministral 3 3B?

Artificial Analysis benchmarking records approximately 201 output tokens per second with a time to first token of about 456 ms. Actual figures vary by provider, region, prompt length, and load, so treat these as a reference point rather than a guarantee.

How does Ministral 3 3B compare to larger Mistral models?

Ministral is Mistral's naming for its compact models, so Ministral 3 3B sits below the larger Mistral and Magistral models in scale. Expect lower cost per token and faster responses, with weaker performance on tasks that require extended reasoning, complex code generation, or heavy long-context work. Many teams route routine traffic to a model like this and escalate harder requests to a larger one.

What context window does Ministral 3 3B support?

We do not currently have a confirmed context window recorded for this model in our database. Check the model card published by the provider you plan to use, since served context limits can also differ between endpoints hosting the same weights.