Mistral Small 4
Mistral Small 4 is a chat-oriented large language model from Mistral, positioned in the company's Small line of efficiency-focused models.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.150 | $0.600 | $0.150 | |
| $0.150 | $0.600 | - |
Prices updated daily. Last check: Oct 11, 2026
Compare API pricing for every Mistral model →Mistral Small 4 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond57.1%
- Humanity's Last Exam3.8%
Agentic & Tool Use
- Terminal-Bench Hard10.6%
- τ²-bench18.4%
Instruction & Long Context
- IFBench32.8%
- Long-Context Reasoning28.3%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Mistral
- Modalities
- Text
Capabilities
- Open Source
- Yes
Strengths & Limitations
Strengths
- Measured output speed of roughly 166 tokens per second in Artificial Analysis testing, suitable for streaming chat interfaces
- Time to first token around 567 ms, keeping perceived response lag low in interactive use
- Positioned in Mistral's Small tier, which targets cost- and latency-sensitive deployments rather than maximum capability
- Part of an established model line, so prompting patterns carried over from earlier Mistral Small generations generally transfer
- Available through multiple inference providers, letting you compare rates and regions in the pricing table on this page
- Small-tier sizing makes it a practical default for high-request-volume pipelines such as extraction and classification
Limitations
- We do not track a confirmed context window for this model — verify the limit with your chosen provider before designing long-document workflows
- Modality support (for example image input) is not confirmed in our data; treat it as text chat unless a provider documents otherwise
- No standardized reasoning, coding, or knowledge benchmark scores are recorded for this entry, making capability comparison against peers harder
- Small-tier models generally trail larger models in the same family on complex multi-step reasoning and long-horizon agentic tasks
- Reported throughput and latency depend heavily on the serving provider and load, so measured figures may not match your deployment
Key Features
About Mistral Small 4
Common Use Cases
Mistral Small 4 fits workloads where per-request latency and volume economics dominate: customer-facing chat and support assistants, retrieval-augmented question answering over pre-filtered context, document and email summarization, structured field extraction, intent and sentiment classification, and content tagging at scale. The sub-second time to first token makes it a reasonable choice for streaming UI experiences where users notice delay before the first word appears, while the ~166 tokens/second generation rate keeps longer responses from feeling slow. For batch pipelines, the Small tier's cost profile usually matters more than its ceiling on hard problems — so it is a common pick for the high-frequency, moderate-difficulty stages of a pipeline, with a larger model reserved for the subset of requests that need deeper reasoning. Before committing it to long-document work, confirm the context limit with your provider, since we do not have that figure verified.
Frequently Asked Questions
How much does Mistral Small 4 cost to use?
Pricing differs by inference provider and by pricing model — some bill separately for input and output tokens, others offer batch or committed-capacity rates. Because rates change frequently, check the pricing table on this page for the current per-provider figures rather than relying on a fixed number.
What is Mistral Small 4 best used for?
It suits chat assistants, summarization, extraction, and classification workloads where fast first-token response and high request volume matter more than maximum reasoning depth. Its measured ~567 ms time to first token and ~166 tokens/second output rate make it a practical fit for streaming interfaces and batch text-processing pipelines.
How fast is Mistral Small 4?
Artificial Analysis measurements put it at roughly 166 output tokens per second with a time to first token of about 567 ms. Actual figures depend on the serving provider, hardware, prompt length, and concurrent load, so treat these as a reference point and benchmark your own traffic.
How does Mistral Small 4 compare to Mistral's larger models?
As a Small-tier model, it is built for lower latency and cost rather than for the hardest reasoning and agentic tasks, where Mistral's larger models generally perform better. A common pattern is to route the bulk of routine requests to a Small-tier model and escalate only the difficult ones.
What context window does Mistral Small 4 support?
We do not have a confirmed context window for this model in our database. Providers sometimes serve the same model with different maximum context settings, so check the documentation of the specific provider you select in the pricing table above.