Skip to main content
MiniMax

MiniMax-M2

MiniMax-M2 is a chat-oriented large language model from MiniMax, offered through multiple inference providers and compared here on hosted API pricing.

Input from
$0.150 / 1M tokens
across 3 providers

API Pricing

Cheapest on Amazon AWS 40% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.150$0.600-
$0.255$1.02-
$0.300$1.20-
$0.300$1.20$0.030

Prices updated daily. Last check: Sep 3, 2026

MiniMax-M2 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
28.9 / 100
Math
78.3 / 100

Reasoning & Knowledge

  • MMLU-Pro82.0%
  • GPQA Diamond77.7%
  • Humanity's Last Exam13.7%

Coding

  • LiveCodeBench82.6%
  • SciCode36.1%

Math

  • AIME 202578.3%

Agentic & Tool Use

  • Terminal-Bench Hard25.8%
  • τ²-bench86.8%

Instruction & Long Context

  • IFBench72.3%
  • Long-Context Reasoning65.0%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
MiniMax
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Available through multiple inference providers, allowing price and latency comparison for the same model
  • Chat-tuned for multi-turn instruction following rather than raw completion
  • Adds a non-US-lab option to model evaluations, useful for vendor diversification
  • Multi-provider hosting reduces single-vendor dependency for production routing
  • Positioned as a cost-oriented alternative in provider catalogs, making it a candidate for high-volume text workloads

Limitations

  • Our metadata does not include a confirmed context window length, so limits must be checked per provider
  • Published throughput and time-to-first-token measurements are not populated in our benchmark record
  • Modality support beyond text is not tracked in our data and should be verified with the provider
  • Standardized public benchmark scores for this entry are not recorded in our database
  • Provider-level features such as tool calling, JSON mode, and structured output may vary between hosts

Key Features

Chat-tuned large language model from MiniMax
Multi-turn conversational prompting and instruction following
Served via hosted inference APIs from multiple providers
OpenAI-compatible chat completion endpoints on most hosting providers
Provider-dependent context window configuration
Part of MiniMax's M-series model line
Listed in the comparison table with live per-provider pricing

About MiniMax-M2

MiniMax-M2 is a text chat model developed by MiniMax, a Chinese AI lab that publishes the MiniMax family of language models. It sits in the M-series line of MiniMax releases and is served through third-party inference APIs, which is how it appears in this catalog. Our database classifies it as a chat model, meaning it is designed for multi-turn conversational prompting and instruction following rather than for embedding, ranking, or media generation. Beyond the model type and creator, our metadata for MiniMax-M2 is limited: we do not currently track a confirmed context window length, modality list, or parameter count for this entry, and the throughput and time-to-first-token figures we have from Artificial Analysis are not populated with usable values. Readers who need exact context limits, supported input types, or tool-calling behavior should confirm those details with the specific provider they plan to route through, since hosted deployments of the same model can differ in configured context length, rate limits, and feature exposure. In practice, MiniMax-M2 is of interest mainly to teams evaluating models outside the largest US labs, where per-token costs and provider availability can differ substantially from more widely deployed alternatives. Because several independent providers can serve the same MiniMax-M2 weights, the pricing table on this page is the useful comparison surface: the model output is nominally the same across hosts, while price, latency, and context configuration vary by provider.

Common Use Cases

MiniMax-M2 fits general text chat workloads: assistant interfaces, drafting and rewriting, summarization, question answering over supplied text, and scripted multi-turn agents where the calling application controls prompt structure. Because it is distributed across several inference providers, it is a reasonable candidate for teams doing cost-sensitive bulk text processing who want to A/B a model against their incumbent without committing to a single vendor. For workloads that depend on a specific long-context limit, image input, or guaranteed structured-output support, confirm those capabilities with the provider first, since our metadata does not record them for this entry.

Frequently Asked Questions

How much does MiniMax-M2 cost to use?

Pricing depends on which inference provider you route through and on the pricing type — input tokens, output tokens, and any cached or batch rates are billed separately and change over time. Check the pricing table on this page for current per-provider rates.

What is MiniMax-M2 best used for?

It is a chat model, so it suits conversational assistants, summarization, drafting, question answering, and multi-turn agent loops driven by your own application logic. It is not an embedding, reranking, or media generation model.

Who created MiniMax-M2?

MiniMax, an AI lab that develops the MiniMax family of language models. MiniMax-M2 belongs to its M-series line.

What context window does MiniMax-M2 support?

We do not have a confirmed context window recorded for this entry. Hosted deployments can be configured with different maximum context lengths, so check the documentation of the specific provider listed in the pricing table.

Does MiniMax-M2 support images or tool calling?

Our database does not track modality or tool-calling details for this model, so we cannot confirm either way. Most providers expose an OpenAI-compatible chat endpoint; verify feature support with the provider you intend to use.

Why do prices differ between providers for the same model?

Providers host the same model on different hardware, with different batching, context configurations, and margin structures. That produces meaningful variation in per-token price and latency, which is why the table on this page lists each provider separately.