Citation and evidence

Xiaomi MiMo-V2.5

11 min full readUpdated 20 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI ModelsChinese AIMixture of ExpertsOpen Source AI

Cite this article

Compare and use this model

Xiaomi MiMo-V2.5 has an article. Its structured comparison facts are awaiting review. Browse the reviewed catalog or suggest a sourced update.

Xiaomi MiMo-V2.5 is an open-weights model family released by Xiaomi in April 2026, made up of two siblings that share a name but solve different problems. One of them, the plain MiMo-V2.5, is a native omnimodal system that takes in text, images, video, and audio. The other, MiMo-V2.5-Pro, is a much larger text-only reasoning and coding model that Xiaomi positions against frontier agents like Claude Opus and GPT-5.4. Both are built on a mixture of experts architecture, handle context windows of up to one million tokens, and ship under the permissive MIT license with weights on Hugging Face.[1][2][3] Xiaomi announced MiMo-V2.5 on 22 April 2026, and the weights of both models appeared on Hugging Face on 27 April.[3][14] In September 2026 the family was succeeded by MiMo-V2.6, whose Pro model keeps the same parameter counts but is omni-modal.[15]

The two are easy to confuse. The "Pro" suffix might suggest the multimodal flagship, but here it is the opposite: Pro is the text specialist, and the smaller, cheaper standard model is the one that can see and hear. Some early coverage blurred the distinction, and at least one downstream tool catalog tagged Pro as accepting image attachments when it does not.[4]

The MiMo line

MiMo is Xiaomi's house brand for large models, and it has moved quickly. The first release, MiMo-7B, arrived in April 2025 as a reasoning-first 7-billion-parameter model that Xiaomi claimed beat OpenAI's o1-mini and Alibaba's QwQ-32B-Preview on math and competitive-programming benchmarks despite its small size.[5] After that the family branched out: MiMo-VL-7B added vision-language understanding, MiMo-Audio-7B handled speech, MiMo-Embodied (November 2025) targeted robotics and embodied tasks, and MiMo-V2-Flash (December 2025) was the first big MoE flagship, a 309B-parameter model aimed at agentic work and also released under MIT.[5][6]

In March 2026 Xiaomi also launched MiMo-V2-Pro, a model with more than 1 trillion total and 42 billion active parameters, and the omni-modal MiMo-V2-Omni through its API; Xiaomi describes MiMo-V2.5-Pro as the successor to MiMo-V2-Pro.[1][19] So by the time V2.5 showed up, Xiaomi had already shipped separate models for reasoning, vision, audio, and embodiment. The V2.5 release is partly an attempt to fold some of that work back together, at least on the standard model, where a single network sees, hears, and acts. Xiaomi's own tagline for it is "a single model that sees, hears, and acts on what it perceives."[3]

The two variants

The cleanest way to think about the pair is reasoning depth versus sensory breadth. Pro goes deep on text and code; the standard model goes wide across modalities. They also differ a lot in size and cost.

MiMo-V2.5 (standard)MiMo-V2.5-Pro
Total parameters310B1.02T
Active parameters per token15B42B
ArchitectureSparse MoEMoE
Routed experts (top-k)256 (top-8)384 (top-8)
Layers48 (1 dense + 47 MoE)70 (1 dense + 69 MoE)
AttentionHybrid SWA + global, 5:1Hybrid SWA + global, 6:1
Context lengthup to 1M tokensup to 1M tokens
Pre-training tokens~48T~27T
PrecisionFP8 (E4M3) mixedFP8 (E4M3) mixed
Modalitiestext, image, video, audiotext only
LicenseMITMIT

Expanded article table

Sources: Xiaomi MiMo site and the two Hugging Face model cards.[1][2][7]

The training-token figures look backwards at first glance, since the smaller model trained on more tokens than the larger one. That is genuinely what the model cards report: roughly 48 trillion tokens for the 310B omnimodal model and roughly 27 trillion for the 1.02T Pro.[2][7] Multimodal pre-training tends to burn through a lot of tokens once you start counting image and audio data, so the gap is not as strange as it sounds, but it is worth flagging because the two figures are easy to mix up.

MiMo-V2.5-Pro: the text-only flagship

Pro is the headline-grabber. It is a 1.02-trillion-parameter MoE language model with 42 billion parameters active on any given token, spread across 384 routed experts with the top 8 selected per token.[7] The network runs 70 layers (one dense layer followed by 69 MoE layers), uses 128 attention heads with 8 key-value heads under grouped-query attention, and ships natively in FP8 (E4M3) weights so it can be served without a separate quantization step.[7]

The attention scheme is the interesting part. Pro interleaves local sliding-window attention with full global attention at a 6:1 ratio, meaning for every six layers that only look at a 128-token window, one layer attends across the whole sequence.[7] Of the 70 layers, 60 use sliding-window attention and 10 use full attention. Xiaomi reports this design, paired with a learnable attention-sink bias, cuts the KV cache footprint by close to seven times compared with a full-attention model of the same size, which is what makes the million-token context window practical to serve.[1][7] The model also carries a three-layer Multi-Token Prediction head for speculative decoding, which Xiaomi says roughly triples output throughput.[7]

Despite the "Pro" badge, this model has no vision or audio encoders. The Hugging Face card describes it plainly as "a Mixture-of-Experts (MoE) language model," and independent coverage confirms it is text-only, built for coding, software engineering, and long-horizon autonomous agents rather than perception.[4][8][9]

MiMo-V2.5: the omnimodal sibling

The standard model is the one with senses. It is a 310B-parameter sparse MoE with 15B active parameters, 256 routed experts (top-8), and 48 layers (one dense plus 47 MoE), using the same hybrid attention idea as Pro but at a 5:1 sliding-window-to-global ratio.[2] Xiaomi describes it as "a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture."[2]

What makes that possible is a pair of dedicated encoders bolted onto the language backbone. There is a 729-million-parameter Vision Transformer with hybrid window attention for images and video, and a 261-million-parameter audio encoder initialized from the weights of Xiaomi's earlier MiMo-Audio model.[2] Both feed into the main network through lightweight projectors, so a single model and a single API call can handle a photo, a video tutorial, or a recorded meeting without you switching tools.[3] It is also the cheaper and faster of the two, which is an unusual place for the multimodal model to sit. At launch Xiaomi's Token Plan billed it at half Pro's credit rate (1x against 2x) and says it surpasses MiMo-V2-Pro in agentic performance while matching MiMo-V2.5-Pro on everyday coding tasks in its internal MiMo Coding Bench "at half the cost."[3][10]

Benchmark claims

Xiaomi's pitch for both models is frontier-level results at a fraction of the token cost, and the numbers it published lean heavily on coding and agentic evaluations rather than raw knowledge tests. As always with vendor-reported benchmarks, these are self-reported and worth treating as claims rather than settled fact.

For Pro, the most cited figure is SWE-Bench Pro, where Xiaomi reports 57.2%. In the comparison table on Xiaomi's model card, that is slightly behind GPT-5.4 (57.7), Claude Opus 4.6 (57.3), GLM 5.1 (58.4) and Kimi K2.6 (58.6).[7][12] The same table lists 78.9 on SWE-bench Verified and 68.4 on Terminal-Bench 2.0.[7] Other Pro scores in the table are 72.9 on the τ3-bench tool-use benchmark (level with GPT-5.4), 63.8 on the Claw-Eval agentic benchmark (pass^3) and a GDPval-AA Elo of 1581.[7] On Humanity's Last Exam Xiaomi reports 48.0 with tools and 34.0 without, against 58.7 and 42.7 for GPT-5.4.[7][11]

The standard model's benchmarks are mostly multimodal. Xiaomi reports 87.7 on Video-MME, 77.9 on MMMU-Pro, and 81.0 on CharXiv reasoning questions, and says the model "stays level with frontier closed-source models" across image, video and multimodal agentic tasks, matching Gemini 3 Pro on video tasks and Claude Sonnet 4.6 on multimodal agentic work.[2][3][12] In Xiaomi's chart, Gemini 3 Pro scores 88.4 on Video-MME and Claude Sonnet 4.6 scores 23.8 on the Claw-Eval multimodal subset, the same as MiMo-V2.5.[2] On the text-and-agent side it scores 62.3 on the general subset of Claw-Eval and 23.8 on the multimodal subset, which Xiaomi frames as sitting "at the Pareto frontier of performance and efficiency."[2][12] Xiaomi also makes a token-efficiency argument. Its Pro launch page says the model reaches 64% pass^3 on ClawEval with about 70,000 tokens per trajectory, "roughly 40-60% fewer tokens than Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4 at comparable capability levels."[1] FinanceFeeds reported further Xiaomi claims that Pro uses 42% fewer tokens than Kimi K2.6 and that the standard model uses roughly half the tokens of similar systems.[11]

Open weights and availability

Both models are fully open-weight under the MIT license, with weights, tokenizer, and model cards published on Hugging Face under the XiaomiMiMo organization.[1][2][7] According to Xiaomi's developer news page, the series entered public beta on Xiaomi's API on 23 April 2026 and was then open-sourced, including base checkpoints, with day-one adaptation for several Chinese AI chips, AMD ROCm and AWS Trainium2.[14] The Hugging Face repositories were created on 27 April 2026.[2][7] Xiaomi released base checkpoints as well as the post-trained instruct versions; the Pro instruct model went through supervised fine-tuning, large-scale agentic reinforcement learning, and a multi-teacher on-policy distillation stage, while the base model ships with a shorter 256K context that the instruct model extends to 1M.[7] For deployment, the model cards give serving recipes for both SGLang and vLLM, and the models are accessible through Xiaomi's own MiMo API alongside third-party inference providers.[2][7][13]

Xiaomi benchmarked the release against DeepSeek-V4 Pro, Kimi K2.6, GLM 5.1, Claude Opus 4.6, Gemini 3.1 Pro and GPT-5.4. On its own numbers Pro does not lead most rows, but it comes close while being downloadable and, by Xiaomi's account, cheaper to run.[7] The split into a text-only flagship and an omnimodal standard model means two downloads, two cost structures and a naming scheme that is easy to misread.

Later developments

In June 2026 Xiaomi and TileRT launched MiMo-V2.5-Pro-UltraSpeed, a faster serving mode of the Pro model that Xiaomi said passed 1,000 tokens per second of decode speed on a 1-trillion-parameter model. It was first offered for a limited, application-based trial from 9 to 23 June 2026 at three times the price of MiMo-V2.5-Pro. After receiving more than 66,000 applications by 23 June, Xiaomi extended the trial; a later notice said the trial had concluded and that a commercial version would follow.[16][20] A July 2026 paper from Xiaomi describes the inference optimizations used to serve the V2.5 series.[17] Xiaomi also retired the older MiMo-V2 API models on 30 June 2026, routing calls for mimo-v2-pro to mimo-v2.5-pro and calls for mimo-v2-omni and mimo-v2-flash to mimo-v2.5.[18]

Successor: MiMo-V2.6

On 21 and 22 September 2026 Xiaomi released MiMo-V2.6, made up of MiMo-V2.6-Pro and MiMo-V2.6-Flash. MiMo-V2.6-Pro keeps the 1.02T total and 42B active parameters of MiMo-V2.5-Pro but adds vision and audio encoders. The V2.6 series was trained with a large mixed-task reinforcement-learning run.[15] In the V2.6 technical report's comparison table, all figures measured by Xiaomi, MiMo-V2.5-Pro is the baseline for the gains:

BenchmarkMiMo-V2.5-ProMiMo-V2.6-Pro
DeepSWE v1.119.071.9
Terminal Bench 2.165.289.9
Terminal Bench 4.01.534.9
Toolathlon-Verified49.176.9
GDPval-AA 2.1 (Elo)11071673
ExploitBench16.647.9

Expanded article table

Source: MiMo-V2.6 technical report, Table 3.[15]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6Xiaomi, "MiMo-V2.5-Pro," mimo.xiaomi.com. mimo.xiaomi.com/mimo-v2-5-pro
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12XiaomiMiMo, "MiMo-V2.5," Hugging Face. huggingface.co/...MiMo-V2.5
  3. ^1 ^2 ^3 ^4 ^5 ^6Xiaomi, "MiMo-V2.5," mimo.xiaomi.com. mimo.xiaomi.com/mimo-v2-5
  4. ^1 ^2NousResearch hermes-agent, "mimo-2.5-pro is text-only, mimo-2.5 is the omnimodal model," GitHub issue #18884. github.com/...18884
  5. ^1 ^2"Xiaomi MiMo," Wikipedia. en.wikipedia.org/...Xiaomi_MiMo
  6. ^"MiMo-V2: Xiaomi AI Models for Reasoning, Multimodal, Voice & API." mimo-v2.org
  7. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16XiaomiMiMo, "MiMo-V2.5-Pro," Hugging Face. huggingface.co/...MiMo-V2.5-Pro
  8. ^RITS, NYU Shanghai, "Xiaomi Releases MiMo-V2.5-Pro: 1T-Parameter Open MoE Matches Frontier Coding Models." rits.shanghai.nyu.edu/...es-frontier-coding-models
  9. ^Tosea.ai, "How to Use MiMo-V2.5-Pro: Complete Guide to Xiaomi's New 1T MoE Model in 2026." tosea.ai/...mimo-v2-5-pro-complete-guide
  10. ^"Xiaomi releases open-weight MiMo-V2.5 AI model, claims frontier-level agentic capability," GSMArena. gsmarena.com/...evel_agentic_capability-news-72585
  11. ^1 ^2"Xiaomi Introduces MiMo V2.5 Featuring Multimodal AI And Enhanced Performance," FinanceFeeds. financefeeds.com/...mimo-v2-5-featuring-multimodal
  12. ^1 ^2 ^3"Xiaomi Releases MiMo-V2.5-Pro and MiMo-V2.5: Matching Frontier Model Benchmarks at Significantly Lower Token Cost," MarkTechPost. marktechpost.com/...significantly-lower-token-cost
  13. ^"Xiaomi releases MiMo-V2.5 and MiMo-V2.5-Pro open-source AI models with MoE architecture," FoneArena. fonearena.com/...xiaomi-mimo-2-5-features
  14. ^1 ^2Xiaomi MiMo, "Xiaomi MiMo-V2.5 series open-sourced & Orbit 100 trillion token plan launched," MiMo developer news. mimo.mi.com/...v2.5-open-sourced
  15. ^1 ^2 ^3LLM-Core Xiaomi, "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement," technical report, September 2026. huggingface.co/...MiMo_V2_6_technical_report.pdf
  16. ^Xiaomi MiMo, "MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS," 8 June 2026. mimo.xiaomi.com/...mimo-tilert-1000tps
  17. ^Xiaomi MiMo Team et al., "Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit," arXiv:2607.13095, July 2026. arxiv.org/...2607.13095
  18. ^"小米 MiMo-V2 系列模型 6 月 30 日正式下线,Pro 版已自动切换至 V2.5" (Xiaomi MiMo-V2 models retired on 30 June), IT Home, 12 June 2026. ithome.com/...664
  19. ^Xiaomi MiMo, "Xiaomi MiMo-V2-Pro," 18 March 2026. mimo.xiaomi.com/mimo-v2-pro
  20. ^Xiaomi MiMo, "Notice on the Extension of the Limited-Time Experience of MiMo-V2.5-Pro-UltraSpeed," MiMo developer news, updated 8 September 2026. mimo.mi.com/...beta-extended

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

3 revisions · v4 · 2,170 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent verification 2026-09-23 (xg05 V7): model-card values checked; V2-Pro comparison and UltraSpeed trial corrected

Cite this page: AI Wiki. "Xiaomi MiMo-V2.5." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/mimo_v2_5

Suggest edit