Ministral 3 3B
Ministral 3 3B is a small chat model from Mistral, part of the Ministral line of compact models aimed at high-throughput, latency-sensitive text workloads.
API Pricing
Cheapest on Amazon AWS — 33% below avg| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.050 | $0.050 | |
| $0.100 | $0.100 |
Prices updated daily. Last check: Oct 10, 2026
Compare API pricing for every Mistral model →Ministral 3 3B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro52.4%
- GPQA Diamond35.8%
- Humanity's Last Exam5.4%
Coding
- LiveCodeBench24.7%
- SciCode15.3%
Math
- AIME 202522.0%
Agentic & Tool Use
- Terminal-Bench Hard0.0%
- Terminal-Bench v2.10.0%
- τ²-bench24.9%
- τ-bench Banking4.7%
Instruction & Long Context
- IFBench26.8%
- Long-Context Reasoning17.0%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Mistral
- Modalities
- Text
Capabilities
- Open Source
- Yes
Strengths & Limitations
Strengths
- Compact ~3B parameter scale, which generally translates to lower cost per token than larger Mistral models
- Measured output throughput of about 201 tokens per second in Artificial Analysis testing
- Time to first token measured at roughly 456 ms, suitable for interactive chat turns
- Part of Mistral's Ministral small-model line, so it shares prompt conventions with other Mistral models for easy swapping
- Small size makes it a practical default tier in cascaded routing setups where a larger model handles escalations
- Throughput profile fits high-volume batch text jobs such as tagging, extraction, and summarization
Limitations
- A ~3B parameter model will generally trail larger Mistral models on multi-step reasoning, long-form coding, and nuanced instruction following
- We do not currently track a confirmed context window for this model — verify with your provider before long-document use
- Multimodal input support is not tracked in our data; do not assume image or audio handling
- Reasoning-heavy benchmark scores are not available in our dataset, making direct quality comparisons with peers harder
- Provider availability for small Mistral models is typically narrower than for the larger flagship models
Key Features
About Ministral 3 3B
Common Use Cases
Ministral 3 3B fits workloads where request volume and response latency drive the architecture more than peak reasoning quality. Typical fits include intent and sentiment classification, structured field extraction from short documents, message drafting and rewriting, autocomplete and suggestion features, chat assistants with narrow scope, and content routing or triage before a larger model is invoked. Its measured sub-half-second time to first token makes it a reasonable choice for streaming interfaces where perceived responsiveness matters, and its ~201 tokens/second output rate supports batch pipelines that generate large volumes of short completions. For tasks involving long multi-step reasoning chains, complex codebase work, or long-context document analysis, a larger model in the Mistral lineup is usually the better match, with Ministral 3 3B handling the routine share of traffic.
Frequently Asked Questions
How much does Ministral 3 3B cost to use?
Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and change frequently. See the pricing table on this page for current per-provider rates rather than relying on a fixed figure.
What is Ministral 3 3B best used for?
It is best suited to high-volume, latency-sensitive text tasks: classification, extraction, rewriting, short chat turns, and triage steps in a larger pipeline. Its compact ~3B scale and measured ~456 ms time to first token favor throughput and responsiveness over deep multi-step reasoning.
How fast is Ministral 3 3B?
Artificial Analysis benchmarking records approximately 201 output tokens per second with a time to first token of about 456 ms. Actual figures vary by provider, region, prompt length, and load, so treat these as a reference point rather than a guarantee.
How does Ministral 3 3B compare to larger Mistral models?
Ministral is Mistral's naming for its compact models, so Ministral 3 3B sits below the larger Mistral and Magistral models in scale. Expect lower cost per token and faster responses, with weaker performance on tasks that require extended reasoning, complex code generation, or heavy long-context work. Many teams route routine traffic to a model like this and escalate harder requests to a larger one.
What context window does Ministral 3 3B support?
We do not currently have a confirmed context window recorded for this model in our database. Check the model card published by the provider you plan to use, since served context limits can also differ between endpoints hosting the same weights.