MiniMax-M3 is a chat-oriented large language model from MiniMax, listed here with measured serving throughput of roughly 107 output tokens per second.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.200 | $0.900 | - | |
| $0.280 | $1.10 | $0.056 | |
| $0.300 | $1.20 | - | |
| $0.300 | $1.20 | - | |
| $0.300 | $1.20 | $0.060 | |
| $0.300 | $1.20 | $0.060 | |
| $0.300 | $1.20 | $0.060 | |
| $0.400 | $2.00 | $0.100 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
MiniMax-M3 is a general-purpose chat model, so the natural fit is conversational workloads where response latency matters: customer-facing assistants, in-product chat, drafting and rewriting tools, summarization of pasted text, and back-end steps in an application pipeline that need natural-language output. The measured sub-second time to first token and roughly 107 tokens per second of generation make it a candidate for streaming UIs where users watch text appear rather than waiting for a complete response. Because we do not have confirmed context window, modality, or task-benchmark data for this model, workloads that depend on very long inputs, image or audio understanding, or verified performance on coding and math evaluations should be validated with your own test set and with the provider's documentation before committing.
Pricing varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are billed differently, and providers change rates over time. Check the pricing table on this page for current per-provider rates rather than relying on a fixed figure.
It is catalogued as a chat model, so it suits conversational assistants, instruction-following tasks, drafting, rewriting, and summarization. Its measured latency profile — about 936 ms to first token and roughly 107 output tokens per second — makes it reasonable for streaming, user-facing interfaces.
In measurements from Artificial Analysis, MiniMax-M3 produced about 106.8 output tokens per second with a time to first token of roughly 936 ms. Both numbers depend on the serving provider, hardware, region, and concurrent load, so treat them as a reference point rather than a guarantee.
We do not have a confirmed context window for MiniMax-M3 in our database. Because context limits can also differ between hosts of the same model, verify the maximum input length with the specific provider you plan to use.
We do not track confirmed modality or tool-calling details for this model, so we make no claim either way. Consult the API documentation of the provider serving MiniMax-M3 for the exact input types and function-calling features it exposes.
MiniMax-M3 was developed by MiniMax, which has released multiple generations of general-purpose language models under the MiniMax name. This entry covers the M3 chat model specifically.