MiMo-V2.5 is a large language model from Xiaomi, part of the company's MiMo family of models served through third-party inference APIs.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.119 | $0.238 | $0.0024 | |
| $0.140 | $0.280 | $0.0028 | |
| $0.140 | $0.280 | $0.0028 | |
| $0.168 | $0.336 | $0.0034 | |
| $0.202 | $0.644 | $0.101 |
Prices updated daily. Last check: Sep 20, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
MiMo-V2.5 is best evaluated for batch and asynchronous text workloads where a two-second wait before the first token is acceptable — content drafting, summarization, data extraction from moderate-length inputs, and offline classification or tagging jobs. Its measured throughput of roughly 54.7 tokens per second is workable for streamed chat responses but the long time to first token makes it a weaker fit for latency-critical paths such as voice agents, autocomplete, or multi-step agent loops where each hop pays that startup cost. Because we do not have confirmed context-window or benchmark data for MiMo-V2.5, teams considering it for coding, math, or long-context retrieval work should run their own evaluation against alternatives listed at similar prices in the table above rather than relying on published capability claims.
Pricing for MiMo-V2.5 varies by inference provider and by pricing type — input tokens, output tokens, and any cached or batch rates are typically billed separately. Because providers change rates and add or drop models frequently, check the pricing table on this page for current per-provider figures.
It suits general text generation work where startup latency is not critical: drafting and rewriting content, summarization, structured extraction, and batch classification. Its measured first-token latency of around 2.3 seconds makes it less suitable for voice interfaces, autocomplete, or agent chains with many sequential model calls.
MiMo-V2.5 is developed by Xiaomi as part of its MiMo model family. It is accessed through third-party inference providers, which are the entities listed in the pricing table on this page.
Independent testing by Artificial Analysis measured roughly 54.7 output tokens per second with a time to first token of about 2,315 ms. Both figures depend on the serving provider, prompt length, and current load, so your observed numbers may differ.
We do not currently have a confirmed context window for MiMo-V2.5 in our database. If your workload involves long documents or large codebases, verify the supported context length directly with the provider you plan to use.
Because we do not track published benchmark scores for MiMo-V2.5, the practical comparison points are its measured speed characteristics and its per-provider price relative to other models in the table above. For capability comparisons, running your own evaluation on representative prompts is the most reliable approach.