MiMo-V2.5-Pro
Xiaomi · released Apr 27, 2026 · XiaomiMiMo/MiMo-V2.5-Pro
- Type
- Open weightsMIT License
- Params
- 1T
- Context
- 1.1M
active per word not recorded by us · about 788K words of context
Our take
Written Sep 2, 2026MiMo-V2.5-Pro is a one-trillion-parameter text model from Xiaomi with a permissive MIT licence and a one-million-token request limit. Its coding score is its standout result, while its agent-task scores sit below zero across every measured subcategory.
Pick this for coding-heavy workloads where its Arena Coding score is over fifty points above its overall mark, or for long-context text work at one million tokens with a genuinely permissive licence. Use it if you want provider choice: ten offers create real price competition. Skip it if you need reliable agent or tool-use behaviour, or if creative writing quality matters — that is its weakest text category.
The case for it
- One-million-token request limit, rare at this scale.
- Coding is its standout category, with an Arena Coding score over fifty points above its overall mark.
- MIT licence permits commercial use, modification and redistribution.
- Ten offers across eight providers, with the cheapest output tier under a third of the most expensive.
The case against it
- Agent-task scores are negative across every measured subcategory, including tool use and recovery.
- Creative writing is its weakest text category, the only one below its overall score.
- Active parameter count is undisclosed, so per-token efficiency is unverified in our data.
How good is it?
An open text model for everyday questions, writing and coding, though it struggles with multi-step agent work and tool calls.
- getting answers to everyday questionsArena Text (overall) · 31st of 168
- drafts, rewrites and editingArena Creative Writing · 40th of 168
- writing and completing codeArena Coding · 24th of 168
- multi-step work it carries out for youArena Agent · 42nd of 55
- calling tools to carry out requestsArena Agent · Tool use · 46th of 55
- getting back on track after a step failsArena Agent · Recovery · 42nd of 55
EverydayGeneral questions and everyday reasoning
Arena Text (overall)31st of 168 · 1467
CodingWriting and fixing code on its own
Arena Coding24th of 168 · 1522
AgenticPlanning, calling tools, staying on task
Arena Agent42nd of 55 · −0.064
Arena Agent is the only board that has scored it for this.
WritingDrafting and rewriting prose
Arena Creative Writing40th of 168 · 1435
Arena Creative Writing is the only board that has scored it for this.
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model12 scoresEvery figure we hold, from 12 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 3 hours and 9 days ago — each listing carries its own date.
Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.
The only listing at 1.1M of context — the other 7 in the table below are not like-for-like. One cheaper row there is outside that comparison: a different quantisation.
- per 1M tokens
- $0.43 in / $0.87 out
- Context served
- 1.1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| GMICloudbf16Through OpenRouter | $0.30 / $0.61checked 15 hours ago | 1.1M945K max reply | 25 tok/s | No | Yesunknown period | Unknown |
| Xiaomifp8Through OpenRouter | $0.43 / $0.87checked 3 hours ago | 1M131K max reply | 30 tok/s | No | Yes30 days | Unknown |
| AtlasCloudfp8Through OpenRouter | $0.43 / $0.87checked 9 hours ago | 1M131K max reply | 33 tok/s | No | Yesunknown period | Unknown |
| OpenRouterOpenRouter's own listing | $0.43 / $0.87checked 3 hours ago | 1.1M | not measured | Unknown | Unknown | Unknown |
| Novita AIDirect and through OpenRouter | $0.52 / $1.04directchecked 3 hours ago$0.48 / $0.96through OpenRouterchecked 3 hours ago | 1M131K max reply through OpenRouter | 27 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| StreamLakeThrough OpenRouter | $0.52 / $1.04checked 3 hours ago | 1M128K max reply | 29 tok/s | No | Yesunknown period | Unknown |
| DigitalOcean GradientThrough OpenRouter | $0.48 / $1.80checked 3 hours ago | 262K236K max reply | 38 tok/s | No | No | Confirmed |
| DeepInfrafp8Direct | $1.00 / $3.00checked 9 days ago | 1M | not measured | Unknown | Unknown | Unknown |
Across the 8 listings we hold: 6 say they do not train on prompts (1 of them only through OpenRouter), 0 say they do and 2 do not say. 2 appear in the zero-retention registry we check (1 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| GMICloudbf16Through OpenRouter | ✗ | ✓ | ✗ |
| Xiaomifp8Through OpenRouter | ✓ | ✓ | ✗ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✗ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Novita AIDirect and through OpenRouter | ✓ | ✓ | ✗ |
| StreamLakeThrough OpenRouter | ✓ | ✓ | ✗ |
| DigitalOcean GradientThrough OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct |
Tool calling: 6 of 8 listings say yes, 1 says no, 1 publishes no parameter list. JSON output: 7 of 8 listings say yes, 1 publishes no parameter list. Strict schema: 2 of 8 listings say yes, 5 say no, 1 publishes no parameter list.
Models people weigh against MiMo-V2.5-Pro
When we formed this view
Recent changes
Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- 1 of 8 listings publishes no parameter list, so what its API accepts is unknown to us.
- We do not hold the active parameter count for it, so how much of it runs on any one token is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 8 listings do not say whether they train on prompts, and 1 answers only through OpenRouter, not for its own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- XiaomiMiMo/MiMo-V2.5-Pro
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- xiaomi-mimo-v2-5-pro