Skip to main content
Open SourceMistral

Mistral Small 4

Mistral Small 4 is a chat-oriented large language model from Mistral, positioned in the company's Small line of efficiency-focused models.

License Open Source
Input from
$0.150 / 1M tokens
across 2 providers

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.150$0.600$0.150
$0.150$0.600-

Prices updated daily. Last check: Oct 11, 2026

Compare API pricing for every Mistral model →

Mistral Small 4 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
9.0 / 100
Output Speed
150 t/s
Latency (TTFT)
560ms

Reasoning & Knowledge

  • GPQA Diamond57.1%
  • Humanity's Last Exam3.8%

Agentic & Tool Use

  • Terminal-Bench Hard10.6%
  • τ²-bench18.4%

Instruction & Long Context

  • IFBench32.8%
  • Long-Context Reasoning28.3%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Mistral
Modalities
Text

Capabilities

Open Source
Yes

Strengths & Limitations

Strengths

  • Measured output speed of roughly 166 tokens per second in Artificial Analysis testing, suitable for streaming chat interfaces
  • Time to first token around 567 ms, keeping perceived response lag low in interactive use
  • Positioned in Mistral's Small tier, which targets cost- and latency-sensitive deployments rather than maximum capability
  • Part of an established model line, so prompting patterns carried over from earlier Mistral Small generations generally transfer
  • Available through multiple inference providers, letting you compare rates and regions in the pricing table on this page
  • Small-tier sizing makes it a practical default for high-request-volume pipelines such as extraction and classification

Limitations

  • We do not track a confirmed context window for this model — verify the limit with your chosen provider before designing long-document workflows
  • Modality support (for example image input) is not confirmed in our data; treat it as text chat unless a provider documents otherwise
  • No standardized reasoning, coding, or knowledge benchmark scores are recorded for this entry, making capability comparison against peers harder
  • Small-tier models generally trail larger models in the same family on complex multi-step reasoning and long-horizon agentic tasks
  • Reported throughput and latency depend heavily on the serving provider and load, so measured figures may not match your deployment

Key Features

•Text chat and instruction-following completions
•Approximately 166 output tokens per second measured throughput (Artificial Analysis)
•Approximately 567 ms time to first token (Artificial Analysis)
•Streaming token output for interactive applications
•Part of Mistral's Small model tier for efficiency-oriented workloads
•Fourth-generation release within the Mistral Small line
•Served by multiple inference API providers tracked in our pricing table

About Mistral Small 4

Mistral Small 4 is a text chat model developed by Mistral, the French AI lab behind the Mistral and Mixtral model lines. It belongs to the Small family, the tier Mistral uses for models sized below its larger flagship releases and aimed at workloads where response speed and cost per request matter as much as raw capability. The "4" denotes its generation within that Small line, following earlier Mistral Small releases. In independent measurements collected by Artificial Analysis, Mistral Small 4 produces roughly 166 output tokens per second with a time to first token of about 567 milliseconds. That combination — sub-second latency before the first token and a steady generation rate — is the practical differentiator for this tier: it makes the model viable for interactive chat surfaces and for pipelines that fan out many short completions in parallel. Measured throughput and latency vary by serving provider, hardware, prompt length, and load, so the figures above should be treated as a reference point rather than a guarantee. We do not currently track a confirmed context window, modality list, or benchmark suite scores for Mistral Small 4 in our database, so those attributes are best verified against the documentation of whichever provider you plan to use. In practice, models in Mistral's Small tier are commonly deployed for assistant-style chat, summarization, extraction, and classification work where a larger model's cost or latency would be hard to justify. Providers serving this model and their current rates are listed in the pricing table on this page.

Common Use Cases

Mistral Small 4 fits workloads where per-request latency and volume economics dominate: customer-facing chat and support assistants, retrieval-augmented question answering over pre-filtered context, document and email summarization, structured field extraction, intent and sentiment classification, and content tagging at scale. The sub-second time to first token makes it a reasonable choice for streaming UI experiences where users notice delay before the first word appears, while the ~166 tokens/second generation rate keeps longer responses from feeling slow. For batch pipelines, the Small tier's cost profile usually matters more than its ceiling on hard problems — so it is a common pick for the high-frequency, moderate-difficulty stages of a pipeline, with a larger model reserved for the subset of requests that need deeper reasoning. Before committing it to long-document work, confirm the context limit with your provider, since we do not have that figure verified.

Frequently Asked Questions

How much does Mistral Small 4 cost to use?

Pricing differs by inference provider and by pricing model — some bill separately for input and output tokens, others offer batch or committed-capacity rates. Because rates change frequently, check the pricing table on this page for the current per-provider figures rather than relying on a fixed number.

What is Mistral Small 4 best used for?

It suits chat assistants, summarization, extraction, and classification workloads where fast first-token response and high request volume matter more than maximum reasoning depth. Its measured ~567 ms time to first token and ~166 tokens/second output rate make it a practical fit for streaming interfaces and batch text-processing pipelines.

How fast is Mistral Small 4?

Artificial Analysis measurements put it at roughly 166 output tokens per second with a time to first token of about 567 ms. Actual figures depend on the serving provider, hardware, prompt length, and concurrent load, so treat these as a reference point and benchmark your own traffic.

How does Mistral Small 4 compare to Mistral's larger models?

As a Small-tier model, it is built for lower latency and cost rather than for the hardest reasoning and agentic tasks, where Mistral's larger models generally perform better. A common pattern is to route the bulk of routine requests to a Small-tier model and escalate only the difficult ones.

What context window does Mistral Small 4 support?

We do not have a confirmed context window for this model in our database. Providers sometimes serve the same model with different maximum context settings, so check the documentation of the specific provider you select in the pricing table above.