Skip to main content
Xiaomi

MiMo-V2.5-Pro

MiMo-V2.5-Pro is a chat-oriented large language model from Xiaomi, positioned as the Pro tier of the company's MiMo V2.5 series.

Input from
$0.400 / 1M tokens
across 4 providers

API Pricing

Cheapest on DigitalOcean 32% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.400$1.50$0.080
$0.435$0.870$0.0036
$0.522$1.04$0.0043
$1.00$3.00$0.200

Prices updated daily. Last check: Aug 29, 2026

MiMo-V2.5-Pro pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
42.9 / 100
Coding
60.2 / 100
Output Speed
37.7 t/s
Latency (TTFT)
5.8s

Reasoning & Knowledge

  • GPQA Diamond86.6%
  • Humanity's Last Exam35.7%

Coding

  • SciCode50.2%

Agentic & Tool Use

  • Terminal-Bench Hard43.2%
  • Terminal-Bench v2.165.2%
  • τ²-bench94.2%
  • τ-bench Banking9.9%

Instruction & Long Context

  • IFBench79.9%
  • Long-Context Reasoning77.7%

Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Xiaomi
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Chat-tuned model suitable for instruction-following and multi-turn dialogue out of the box
  • Pro-tier position within Xiaomi's MiMo V2.5 line, above the family's lighter configurations
  • Independently measured throughput data available from Artificial Analysis (≈37.7 output tokens/sec), so performance expectations can be set before deployment
  • Adds a Xiaomi-built option to a provider landscape dominated by a handful of Western and Chinese labs, useful for vendor diversification
  • Steady streaming rate makes long-form generation predictable for drafting and summarisation tasks
  • Listed across providers on this page, so throughput and cost trade-offs can be compared side by side

Limitations

  • Measured time to first token of roughly 5,786 ms makes it poorly suited to latency-sensitive interactive applications
  • Output throughput of ≈37.7 tokens per second is moderate, so long responses take noticeable wall-clock time
  • We do not track a confirmed context window for this model — check the serving provider's documentation before planning long-document workloads
  • Modality support, tool calling, and structured output behaviour are not recorded in our database and should be verified directly
  • Less third-party benchmark coverage and community tooling than more widely deployed model families

Key Features

Chat and instruction-following interface for multi-turn conversations
Pro-tier configuration of the Xiaomi MiMo V2.5 series
Measured output throughput of approximately 37.7 tokens per second
Measured time to first token of approximately 5,786 ms (Artificial Analysis)
Streaming token output for progressive response rendering where the provider supports it
Available through hosted inference providers listed in the pricing table on this page
Text generation for drafting, summarisation, and coding assistance workloads

About MiMo-V2.5-Pro

MiMo-V2.5-Pro is a conversational large language model developed by Xiaomi as part of its MiMo model line. It sits in the Pro tier of the V2.5 generation, which places it above the smaller or lighter configurations that typically accompany a Pro release in the same family. Xiaomi is one of a number of Chinese technology companies that have moved from consumer hardware into publishing their own foundation models, and MiMo is the umbrella name used for that effort. Our database records MiMo-V2.5-Pro as a chat model — that is, a model intended for instruction-following and multi-turn dialogue rather than a specialised embedding, reranking, or media-generation model. Third-party measurements from Artificial Analysis put its output throughput at roughly 37.7 tokens per second with a time to first token of about 5,786 ms. The elevated latency to first token is characteristic of models that perform extended internal processing before emitting a response, though we do not track a confirmed context window, modality list, or tool-calling specification for this entry, so those details should be verified against the serving provider's documentation. In practice, a Pro-tier chat model of this kind is generally deployed for assistant-style workloads, content drafting, summarisation, and code assistance, where response quality matters more than sub-second responsiveness. Because MiMo-V2.5-Pro's measured time to first token is on the slower side and its streaming rate is moderate, it is a better fit for asynchronous or batch-style interactions than for latency-critical user-facing turns. Availability and serving characteristics vary by provider — see the pricing table on this page for who currently hosts it.

Common Use Cases

MiMo-V2.5-Pro fits general assistant workloads where quality of the completed answer matters more than how quickly the first token appears: long-form content drafting, document and transcript summarisation, question answering over supplied text, and code explanation or generation. Its measured latency profile — a multi-second wait before generation begins, followed by a steady mid-range streaming rate — points toward background jobs, batch processing pipelines, and asynchronous agent steps rather than live chat widgets, voice interfaces, or autocomplete. Teams evaluating Chinese-developed model families alongside more common options may also use it as a comparison point or as a secondary provider for redundancy. For high-volume classification or routing, a smaller and faster model is usually the better economic choice.

Frequently Asked Questions

How much does MiMo-V2.5-Pro cost to use?

Pricing depends on which provider hosts the model and on the pricing type — input tokens, output tokens, and any cached or batch rates are typically billed separately, and providers change rates frequently. Check the pricing table on this page for the current per-provider figures.

What is MiMo-V2.5-Pro best used for?

It is a chat model, so it suits instruction-following and dialogue workloads: drafting text, summarising documents, answering questions, and assisting with code. Given its measured multi-second time to first token, it works best in asynchronous or batch settings rather than latency-critical interactive interfaces.

Who makes MiMo-V2.5-Pro?

It is developed by Xiaomi, under its MiMo model line. MiMo-V2.5-Pro is the Pro tier of the V2.5 generation.

How fast is MiMo-V2.5-Pro?

Third-party measurements from Artificial Analysis record approximately 37.7 output tokens per second and a time to first token of roughly 5,786 ms. Actual figures vary by provider, prompt length, and load, so compare the hosts listed on this page.

What context window does MiMo-V2.5-Pro support?

We do not currently track a confirmed context window for this model. Consult the documentation of whichever provider you plan to use, since serving configurations can also cap the usable window below the model's native limit.