MiniMax-M2
MiniMax-M2 is a chat-oriented large language model from MiniMax, offered through multiple inference providers and compared here on hosted API pricing.
API Pricing
Cheapest on Amazon AWS — 40% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.150 | $0.600 | - | |
| $0.255 | $1.02 | - | |
| $0.300 | $1.20 | - | |
| $0.300 | $1.20 | $0.030 |
Prices updated daily. Last check: Sep 3, 2026
MiniMax-M2 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro82.0%
- GPQA Diamond77.7%
- Humanity's Last Exam13.7%
Coding
- LiveCodeBench82.6%
- SciCode36.1%
Math
- AIME 202578.3%
Agentic & Tool Use
- Terminal-Bench Hard25.8%
- τ²-bench86.8%
Instruction & Long Context
- IFBench72.3%
- Long-Context Reasoning65.0%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- MiniMax
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Available through multiple inference providers, allowing price and latency comparison for the same model
- Chat-tuned for multi-turn instruction following rather than raw completion
- Adds a non-US-lab option to model evaluations, useful for vendor diversification
- Multi-provider hosting reduces single-vendor dependency for production routing
- Positioned as a cost-oriented alternative in provider catalogs, making it a candidate for high-volume text workloads
Limitations
- Our metadata does not include a confirmed context window length, so limits must be checked per provider
- Published throughput and time-to-first-token measurements are not populated in our benchmark record
- Modality support beyond text is not tracked in our data and should be verified with the provider
- Standardized public benchmark scores for this entry are not recorded in our database
- Provider-level features such as tool calling, JSON mode, and structured output may vary between hosts
Key Features
About MiniMax-M2
Common Use Cases
MiniMax-M2 fits general text chat workloads: assistant interfaces, drafting and rewriting, summarization, question answering over supplied text, and scripted multi-turn agents where the calling application controls prompt structure. Because it is distributed across several inference providers, it is a reasonable candidate for teams doing cost-sensitive bulk text processing who want to A/B a model against their incumbent without committing to a single vendor. For workloads that depend on a specific long-context limit, image input, or guaranteed structured-output support, confirm those capabilities with the provider first, since our metadata does not record them for this entry.
Frequently Asked Questions
How much does MiniMax-M2 cost to use?
Pricing depends on which inference provider you route through and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and change over time. Check the pricing table on this page for current per-provider rates.
What is MiniMax-M2 best used for?
It is a chat model, so it suits conversational assistants, summarization, drafting, question answering, and multi-turn agent loops driven by your own application logic. It is not an embedding, reranking, or media generation model.
Who created MiniMax-M2?
MiniMax, an AI lab that develops the MiniMax family of language models. MiniMax-M2 belongs to its M-series line.
What context window does MiniMax-M2 support?
We do not have a confirmed context window recorded for this entry. Hosted deployments can be configured with different maximum context lengths, so check the documentation of the specific provider listed in the pricing table.
Does MiniMax-M2 support images or tool calling?
Our database does not track modality or tool-calling details for this model, so we cannot confirm either way. Most providers expose an OpenAI-compatible chat endpoint; verify feature support with the provider you intend to use.
Why do prices differ between providers for the same model?
Providers host the same model on different hardware, with different batching, context configurations, and margin structures. That produces meaningful variation in per-token price and latency, which is why the table on this page lists each provider separately.