MiMo-V2.6-Pro is a chat-oriented large language model from Xiaomi, part of the company's MiMo model line, with measured output of roughly 121 tokens per second.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.435 | $0.870 | $0.0036 |
Prices updated daily. Last check: Sep 22, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
MiMo-V2.6-Pro fits general assistant workloads where response streaming speed matters: customer-facing chat, in-product help and Q&A, drafting and rewriting text, summarization, and conversational front-ends where a roughly 1.2-second initial delay followed by ~121 tokens per second of output feels responsive. As the Pro variant of its generation, it targets tasks that need more capability than a lightweight chat model while still running at interactive speeds. Teams building agent loops or batch pipelines should first confirm the context window and any tool-calling support with their chosen provider, since those specifics are not tracked here. It is also a reasonable candidate for teams doing multi-vendor evaluations who want to price a Xiaomi-built option alongside more common model families.
Pricing depends on the provider hosting the model and on the pricing type — per-million input tokens, per-million output tokens, and any cached-input or batch rates are all billed separately and change frequently. Check the pricing table on this page for current rates from each tracked provider rather than relying on a figure quoted elsewhere.
It is a chat model, so it suits conversational assistants, drafting and editing, summarization, and Q&A interfaces. Its measured throughput of about 121 output tokens per second and roughly 1.2-second time to first token make it a fit for interactive, user-facing applications where responses stream as they generate.
Xiaomi. MiMo is the company's series of large language models, and V2.6-Pro is the Pro tier within the V2.6 generation of that series.
Artificial Analysis measured approximately 120.97 output tokens per second with a time to first token of about 1,228 milliseconds. Actual figures vary by provider, region, prompt length, and server load, so treat these as a reference point rather than a guarantee.
We do not have confirmed multimodal support for this model in our database. If your workload requires image or audio input, verify against the documentation of the specific provider you plan to use before committing.
Without tracked benchmark suite results, the most concrete comparison points we have are serving metrics — roughly 121 output tokens per second and about 1.2 seconds to first token — plus whatever per-token price the providers in the table above charge. For quality comparisons, run your own evaluation on representative prompts from your workload.