MiMo-V2.5-Pro
MiMo-V2.5-Pro is a chat-oriented large language model from Xiaomi, positioned as the Pro tier of the company's MiMo V2.5 series.
API Pricing
Cheapest on DigitalOcean — 32% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.400 | $1.50 | $0.080 | |
| $0.435 | $0.870 | $0.0036 | |
| $0.522 | $1.04 | $0.0043 | |
| $1.00 | $3.00 | $0.200 |
Prices updated daily. Last check: Aug 29, 2026
MiMo-V2.5-Pro pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond86.6%
- Humanity's Last Exam35.7%
Coding
- SciCode50.2%
Agentic & Tool Use
- Terminal-Bench Hard43.2%
- Terminal-Bench v2.165.2%
- τ²-bench94.2%
- τ-bench Banking9.9%
Instruction & Long Context
- IFBench79.9%
- Long-Context Reasoning77.7%
Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Xiaomi
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Chat-tuned model suitable for instruction-following and multi-turn dialogue out of the box
- Pro-tier position within Xiaomi's MiMo V2.5 line, above the family's lighter configurations
- Independently measured throughput data available from Artificial Analysis (≈37.7 output tokens/sec), so performance expectations can be set before deployment
- Adds a Xiaomi-built option to a provider landscape dominated by a handful of Western and Chinese labs, useful for vendor diversification
- Steady streaming rate makes long-form generation predictable for drafting and summarisation tasks
- Listed across providers on this page, so throughput and cost trade-offs can be compared side by side
Limitations
- Measured time to first token of roughly 5,786 ms makes it poorly suited to latency-sensitive interactive applications
- Output throughput of ≈37.7 tokens per second is moderate, so long responses take noticeable wall-clock time
- We do not track a confirmed context window for this model — check the serving provider's documentation before planning long-document workloads
- Modality support, tool calling, and structured output behaviour are not recorded in our database and should be verified directly
- Less third-party benchmark coverage and community tooling than more widely deployed model families
Key Features
About MiMo-V2.5-Pro
Common Use Cases
MiMo-V2.5-Pro fits general assistant workloads where quality of the completed answer matters more than how quickly the first token appears: long-form content drafting, document and transcript summarisation, question answering over supplied text, and code explanation or generation. Its measured latency profile — a multi-second wait before generation begins, followed by a steady mid-range streaming rate — points toward background jobs, batch processing pipelines, and asynchronous agent steps rather than live chat widgets, voice interfaces, or autocomplete. Teams evaluating Chinese-developed model families alongside more common options may also use it as a comparison point or as a secondary provider for redundancy. For high-volume classification or routing, a smaller and faster model is usually the better economic choice.
Frequently Asked Questions
How much does MiMo-V2.5-Pro cost to use?
Pricing depends on which provider hosts the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are typically billed separately, and providers change rates frequently. Check the pricing table on this page for the current per-provider figures.
What is MiMo-V2.5-Pro best used for?
It is a chat model, so it suits instruction-following and dialogue workloads: drafting text, summarising documents, answering questions, and assisting with code. Given its measured multi-second time to first token, it works best in asynchronous or batch settings rather than latency-critical interactive interfaces.
Who makes MiMo-V2.5-Pro?
It is developed by Xiaomi, under its MiMo model line. MiMo-V2.5-Pro is the Pro tier of the V2.5 generation.
How fast is MiMo-V2.5-Pro?
Third-party measurements from Artificial Analysis record approximately 37.7 output tokens per second and a time to first token of roughly 5,786 ms. Actual figures vary by provider, prompt length, and load, so compare the hosts listed on this page.
What context window does MiMo-V2.5-Pro support?
We do not currently track a confirmed context window for this model. Consult the documentation of whichever provider you plan to use, since serving configurations can also cap the usable window below the model's native limit.