Ministral 3 3B is a small chat model from Mistral, part of the Ministral line of compact models aimed at high-throughput, latency-sensitive text workloads.
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.050 | $0.050 | |
| $0.100 | $0.100 |
Prices updated daily. Last check: Sep 5, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Ministral 3 3B fits workloads where request volume and response latency drive the architecture more than peak reasoning quality. Typical fits include intent and sentiment classification, structured field extraction from short documents, message drafting and rewriting, autocomplete and suggestion features, chat assistants with narrow scope, and content routing or triage before a larger model is invoked. Its measured sub-half-second time to first token makes it a reasonable choice for streaming interfaces where perceived responsiveness matters, and its ~201 tokens/second output rate supports batch pipelines that generate large volumes of short completions. For tasks involving long multi-step reasoning chains, complex codebase work, or long-context document analysis, a larger model in the Mistral lineup is usually the better match, with Ministral 3 3B handling the routine share of traffic.
Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and change frequently. See the pricing table on this page for current per-provider rates rather than relying on a fixed figure.
It is best suited to high-volume, latency-sensitive text tasks: classification, extraction, rewriting, short chat turns, and triage steps in a larger pipeline. Its compact ~3B scale and measured ~456 ms time to first token favor throughput and responsiveness over deep multi-step reasoning.
Artificial Analysis benchmarking records approximately 201 output tokens per second with a time to first token of about 456 ms. Actual figures vary by provider, region, prompt length, and load, so treat these as a reference point rather than a guarantee.
Ministral is Mistral's naming for its compact models, so Ministral 3 3B sits below the larger Mistral and Magistral models in scale. Expect lower cost per token and faster responses, with weaker performance on tasks that require extended reasoning, complex code generation, or heavy long-context work. Many teams route routine traffic to a model like this and escalate harder requests to a larger one.
We do not currently have a confirmed context window recorded for this model in our database. Check the model card published by the provider you plan to use, since served context limits can also differ between endpoints hosting the same weights.