Xiaomi MiMo-V2.6 Explained: MIT-Licensed Open Weights and an Open RL Stack
Most open-weight releases hand you a checkpoint and a benchmark table. MiMo-V2.6, released by Xiaomi on 21 to 22 September 2026 (outlets differ by a day, likely a time-zone effect), hands you the checkpoint, the MIT license, and a good part of the machinery that produced it: more than 7,000 reinforcement learning task environments, a fork of the verl RL framework, and Docker images for reproducing the setup.
That matters because the frontier of open models has moved from pretraining recipes to post-training recipes. Whoever can show how agent-grade reinforcement learning (RL) is done, with real cost figures, changes what a small lab can attempt. This post explains what shipped, how the architecture works, what the RL run cost, which numbers are solid and which are vendor-reported, and what it takes to serve the models yourself.
What this covers: lineage and what changed, the sparse Mixture-of-Experts architecture, the RL training pipeline, benchmarks with caveats, deployment and pricing math, failure modes, and a decision matrix against peer open models.
Context and Background
Xiaomi’s MiMo line is the consumer-electronics maker’s foundation model effort. Per the Wikipedia timeline, MiMo-7B arrived on 30 April 2025 under MIT, MiMo-V2-Flash (309 billion total parameters, 15 billion active) on 17 December 2025, and a proprietary MiMo-V2-Pro (about 1 trillion parameters, 42 billion active) in March 2026. MiMo-V2.5 and V2.5-Pro followed in April 2026, with the Pro weights released under MIT. One outlet dates the V2.5 open-sourcing to 29 June; treat the exact V2.5 dates as loosely sourced.
The V2.6 series, per Xiaomi’s release notes, consists of MiMo-V2.6-Pro, MiMo-V2.6-Flash, and MiMo-V2.6-Pro-UltraSpeed, an inference mode of Pro rather than a separate model. Xiaomi also published MiMo-V2.6-Distill-Qwen-9B, a research checkpoint fine-tuned from Alibaba’s Qwen3.5-9B. It is not a small Pro; it is a vehicle for studying agentic RL at low cost, and its dependence on Qwen means you should read the dependency and license chain before using it commercially.
This is the crowded end of the market. If you have read our coverage of DeepSeek V4, Kimi K3, GLM-5.2 and Qwen 3.6, you know the pattern: trillion-scale sparse MoE, long context, aggressive agent-benchmark claims. For the general mechanics, see our primer on mixture-of-experts LLM architecture.
What separates MiMo-V2.6 is not a new architecture. Pro and Flash inherit the V2 design. The novelty is in the training story: a short, very large, fully asynchronous RL run whose cost was disclosed, and whose environments were released.
What the “open” in open weight means here
Three artifacts carry different licenses, and mixing them up is a common mistake. The Pro and Flash model repositories on Hugging Face list MIT. The RL environment dataset, MiMo-V2.6-RL-oss, lists Apache 2.0. The training-code repository, a fork of verl on a mimo-oss branch, carries Apache-2.0 as inherited from upstream.
All are permissive and commercial-friendly. But “open weights plus environments plus framework fork” is still not “fully reproducible.” Pretraining data and the pretraining code are not part of the release, and the released RL code is a framework fork, not a one-command replica of the six-day run. We return to this in the RL section.
MiMo-V2.6 Architecture: A Sparse MoE With Mostly-Local Attention
Direct answer: MiMo-V2.6-Pro is a 1.02-trillion-parameter sparse Mixture-of-Experts model that activates 42 billion parameters per token, using 70 layers of which 60 are 128-token sliding-window attention and 10 are global attention. Flash is a 309-billion-parameter sibling with 15 billion active. Both are natively multimodal with a 1,048,576-token context.

Figure 1: MiMo-V2.6-Pro data path. Modality encoders feed a 70-layer backbone that mixes local and global attention, routes each token to 8 of 384 experts, and drafts several tokens ahead with a multi-token-prediction head.
Figure 1 summarizes the model card specifications. Numbers below come from the Hugging Face cards for MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL.
| Spec | Pro | Flash |
|---|---|---|
| Total parameters | 1.02 T | 309 B |
| Active per token | 42 B | 15 B |
| Layers | 70 (60 SWA + 10 GA) | 48 (39 SWA + 9 GA) |
| Hidden size | 6,144 | 4,096 |
| Attention heads (Q/KV) | 128/8 (SWA and GA) | 64/8 SWA, 64/4 GA |
| Routed experts | 384, 8 active | 256, 8 active |
| Sliding window | 128 tokens | 128 tokens |
| Context | 1M tokens (1,048,576 on API) | 1M tokens |
| MTP draft module | 5 layers, 7 extra tokens | 5 SWA layers |
The vision encoder is the 681-million-parameter MiMo ViT (28 layers, 24 sliding-window and 4 global). Audio uses a 308-million-parameter AudioTokenizer with 20 residual vector quantization codebooks on the Pro card, plus a 127-million-parameter audio patch encoder. Xiaomi’s line for Pro is text, image, video, and audio in one model.
Why 128-token windows are the interesting choice
A 128-token sliding window is extremely short. In 60 of Pro’s 70 layers, a token can only attend to its previous 128 neighbors. Long-range information travels through the 10 global layers, and through the depth of the network as windows stack.
The payoff is the key-value (KV) cache. In a standard transformer the KV cache grows linearly with context in every layer. Here, only global layers grow. Counting layers alone, Pro has 10 of 70 layers (14%) whose cache scales with context, Flash has 9 of 48 (19%). Sliding-window layers keep a fixed 128-token cache regardless of whether the prompt is 4K or 1M tokens.
As a rough, illustrative estimate that ignores head dimension differences and any cache compression, the context-dependent cache is about 5 to 7 times smaller than an all-global model of the same depth. That is why a million-token window is plausible to serve at all. It is derived from the published layer counts, not a Xiaomi-published figure.
The risk is obvious: retrieval-heavy tasks that need many distant tokens depend on just 10 layers. We cover long-context degradation in the limitations section.
Sparsity math
Pro activates 42 B of 1,020 B parameters: about 4.1%. Flash activates 15 B of 309 B: about 4.9%. At the expert level, Pro routes each token to 8 of 384 experts (2.1%), Flash to 8 of 256 (3.1%). Per-token compute therefore resembles a 42-billion-parameter dense model, while memory footprint resembles a trillion-parameter one.
That asymmetry drives everything about deployment. Compute is cheap per token; memory and interconnect are not. Expert placement across GPUs, all-to-all communication and load balance dominate serving design, a topic we detail in expert-parallel MoE inference serving.
Multi-token prediction as a built-in speculative decoder
Both models ship a multi-token-prediction (MTP) module: on Pro, a 5-layer speculative decoder that predicts 7 subsequent tokens per forward pass with a 1,024-token window. In use, the main model verifies the drafted tokens and accepts the longest correct prefix, which is the standard speculative decoding trade: extra draft compute for fewer sequential decode steps.
Xiaomi’s UltraSpeed mode claims up to 20 times Pro’s output speed at ten times the price. Xiaomi does not publish in the sources I could access how UltraSpeed achieves this, so I will not guess at the mechanism; MTP drafting, dedicated hardware, or batching policy could each contribute.
Training: A Six-Day, Fully Asynchronous RL Run
Xiaomi’s stated thesis is that agent performance scales with three RL ingredients at once: RL compute, the diversity of task environments, and grader compute. The V2.6 release is the evidence it offers, and it is also where the open-source claim has real substance.

Figure 2: The reported V2.6 post-training loop. Prompts from a large environment pool are rolled out asynchronously, graded, ranked, and fed to a GRPO update with the MoE router frozen; the RL checkpoint is then fused with specialist teachers through MOPD2.
What Xiaomi discloses
Per the Pro-RL model card and Xiaomi’s release notes, the recipe is fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches: 1,568 prompts times 16 rollouts per step. Contexts run up to 1M tokens, and each step generates roughly 3.5 to 3.7 billion training tokens. Each of Pro and Flash completed 30 RL steps in under six days, with coding, general-agent, visual, and cybersecurity tasks mixed in one run.
Reported RL compute cost, per Xiaomi as relayed by several outlets: about $2.62 million for Pro and $850,000 for Flash. One outlet headlines the sum, $3.47 million, which matches 2.62 plus 0.85. Reporting states these figures exclude pretraining and other development work.
Working the numbers
These are derived from the disclosed figures; treat them as arithmetic, not Xiaomi statements.
- Rollouts per step: 1,568 times 16 equals 25,088.
- Rollouts per 30-step run: 25,088 times 30 equals 752,640, which lines up with the “approximately 750,000 trajectories” figure. Reports describe that number as cumulative across both models, yet a single 30-step run at this batch size already gives about 753K. Either Flash used a smaller effective batch or the reporting is loose. I could not resolve this from the sources.
- Tokens per rollout: 3.5 to 3.7 billion divided by 25,088 is roughly 140,000 to 147,000 tokens per trajectory on average. These are long agentic episodes, not chat turns.
- Cost per step (Pro): $2.62 million over 30 steps is about $87,000 per step.
- Cost per trajectory (Pro): about $3.50 if the 753K figure applies to Pro alone.
A per-trajectory cost of a few dollars explains why verifier and environment quality matter more than raw volume. At that price, a rollout wasted on a broken task or a gameable grader is real money.
Why frozen routing matters
Xiaomi’s notes say the MoE router is frozen during RL to suppress expert load drift. This is a sensible engineering choice. In sparse MoE models, policy-gradient updates can shift routing distributions between the sampler and the trainer, and with asynchronous rollouts the policy that generated a trajectory is already stale when the update lands. If routing also drifts, the importance ratios GRPO relies on become noisy for reasons unrelated to the task.
Freezing the router trades expressiveness for stability: experts can specialize further, but the assignment of tokens to experts cannot adapt to the new skill distribution. For a 30-step run this is a reasonable bargain. For long-horizon RL programs it may become a ceiling, and whether it does is an open question the release does not answer.
Ranked grading, not just pass/fail
Reporting on the technical report describes ranking successful trajectories to redistribute reward toward higher-quality solutions, and “groupwise agentic grading” that analyzes execution traces. This targets a known weakness of binary verifiers: if 14 of 16 rollouts pass a unit test, the group advantage carries almost no signal about which solutions were cleaner, safer, or cheaper.
Adding a quality ranking among passing samples gives GRPO a gradient inside the success region. The cost is grader compute, the third scaling axis Xiaomi names. A trace-reading grader is itself a model call per trajectory, at up to a million tokens of context.
MOPD2: distilling specialists back into one model
After RL, the recipe uses Multi-Prefix Multi-Teacher On-Policy Distillation, MOPD2. The Flash-MOPD card describes three streams: standard on-policy distillation where RL teachers supervise the student’s own rollouts; teacher-prefix distillation, where the student continues from histories a teacher produced; and SFT-prefix distillation from synthetic demonstration histories. Teachers fall into two groups, mixRL teachers for verifiable tasks and SFT teachers for open-domain tasks lacking reliable rewards.
The stated target is a concrete failure mode: tool-call repetition, where an agent keeps issuing near-identical calls and looks busy while making no progress. Xiaomi’s card shows a response-level repetition rate chart before and after; the excerpt I could read did not include the numbers, so I cannot quote a reduction.
What is actually open in the RL stack
This is the part worth analytical attention, because “RL stack open-sourced” can mean many things. What I could verify:
| Component | Where | License | Notes |
|---|---|---|---|
| RL task environments | MiMo-V2.6-RL-oss dataset | Apache 2.0 | 7,780 rows, 12.1 GB Parquet |
| Training framework | github.com/XiaomiMiMo/verl, branch mimo-oss |
Apache-2.0 | Fork of upstream verl |
| Execution images | Docker Hub xiaomimimo/mimo-v2.6-rl-oss |
Not verified | Per dataset card |
| Weights and technical report | Hugging Face | MIT (weights) | Pro-RL, Flash-RL, Flash-MOPD, Distill-Qwen-9B |
The dataset’s five domains, by row count: code 2,700, web development 2,090, cyber 1,000, music 1,000 (symbolic composition), general knowledge work 989. Rows carry a prompt, an agent_name such as mimo_swe_agent, a reward_model field with ground truth or style criteria, and extra_info including Docker configuration. Verification is domain specific: executable tests, rule checks, visual grading, and rubric judging.
Note the numbers: the headline says “7,000+ environments” while the dataset shows 7,780 rows. That is consistent (7,780 exceeds 7,000), though rows and environments may not map one to one.
What is not evidently released: the pretraining corpus, the pretraining code, the frozen-router implementation details beyond what the fork contains (the fork’s README focuses on upstream verl, and I could not confirm MiMo-specific patches from the page), and the grader models used for ranked and agentic grading. An independent write-up makes the sharpest point: outside testing evaluates the released model, but it “does not independently validate Xiaomi’s claims about which parts of its RL system produced the gains.”
Why this matters beyond Xiaomi
Environments are the scarce resource in agent RL. Building thousands of containerized tasks with reliable verifiers takes far more engineering than running GRPO. Releasing them lets a university lab or a startup run a scaled-down replication on the 9B distill checkpoint, and gives evaluators a substrate for building harder agent benchmarks. That is the durable contribution, whatever the leaderboard does next month.
For teams building agent systems, our notes on AI agent trajectory evaluation patterns apply directly to how you would grade and inspect rollouts like these.
Capabilities and Benchmarks: Read the Provenance of Every Number
Benchmark tables for agent models are where careless writing goes wrong, so this section separates three kinds of numbers: Xiaomi’s before-and-after RL gains, the Hugging Face model-card scores, and third-party indices.

Figure 3: Where each MiMo-V2.6 number comes from. RL deltas and model-card scores are vendor-reported; the Artificial Analysis index is the main independent reference.
What changed versus V2.5
The clearest statement of what V2.6 changed is Xiaomi’s own RL delta. On DeepSWE v1.1, Flash moved from 48.8 to 65.7 (about 17 points) and Pro from 58.4 to 72.6 (about 14 points). Average pass rate on the training task mix rose about 25% for Flash and 12% for Pro. API pricing is reported unchanged from the V2.5 series, so the improvement arrives at constant price. I could not verify whether the “before” checkpoint equals the released V2.5 model or an internal pre-RL checkpoint; the release notes describe RL gains, so read them as within-run deltas.
Note also that the Pro model card lists DeepSWE v1.1 at 71.9, while Xiaomi’s post-RL delta is 72.6 (72.57). The gap is 0.7 points and probably reflects different evaluation runs or the MOPD stage; nothing in the sources explains it. Small discrepancies like this are why you should pin your own evaluation.
Model-card scores
From the Pro and Flash cards (vendor-reported):
| Benchmark | Pro | Flash |
|---|---|---|
| DeepSWE v1.1 | 71.9 | 67.9 |
| Terminal Bench 2.1 | 89.9 | 87.6 |
| Toolathlon-Verified | 76.9 | 73.6 |
| CyberGym | 94.0 | 95.1 |
| OSWorld-Verified | 82.0 | not listed in the excerpt |
| AutomationBench v1.0.6 | 53.1 | not listed |
| ProgramBench | 26.5 | not listed |
| Agents’ Last Exam | 31.6 | not listed |
| Terminal Bench 4.0 | 34.9 | not listed |
| MiMo VisualCoding | 72.3 | not listed |
| ExploitGym / ExploitBench | 17.8 / 47.9 | not listed |
Two things stand out. First, Flash sits within a few points of Pro on several agent tasks and beats it on CyberGym (95.1 against 94.0). That is unusual for a model with a third of the active parameters and suggests RL and task coverage, not raw size, drive these particular scores. Second, the harder, newer tests (ProgramBench at 26.5, Terminal Bench 4.0 at 34.9, ExploitGym at 17.8) sit far below the saturating ones (Terminal Bench 2.1 near 90). When one suite is close to ceiling, differences of one or two points are noise.
Independent and comparative claims
Artificial Analysis scored Pro at 46 (46.32 on Intelligence Index v4.3), which Xiaomi and several outlets describe as the highest among open-weight models. Reporting on comparators varies: one outlet says Pro tied Grok 4.7 at 46 and put DeepSeek V4.1 Pro at 36; another cites GLM-5.3 at 45 and Kimi K3 at 44. Xiaomi’s own release note says Pro surpasses Kimi K3 and Qwen3.8 Max on that index. I cannot reconcile these lists, and index rankings move as models are re-run, so check the live Artificial Analysis page rather than trusting a screenshot.
On closed models, Xiaomi says Pro is on par with Claude Opus 5 and GPT-5.6 Sol on agent benchmarks. Reporting is more careful: on AutomationBench Pro scores 53.1 against 50.3 for Opus 5, on Terminal Bench 2.1 it scores 89.9 against 89.1, and Agents’ Last Exam is a tie. But VentureBeat notes Xiaomi is “not showing an across-the-board victory over the strongest closed models” and that Opus 5 remains ahead on DeepSWE v1.1, ProgramBench, and Terminal Bench 4.0.
Caveats that apply to all of this
- Vendor-reported, mostly. The agent scores come from Xiaomi’s harnesses. Scaffold, tool set, turn limits, and timeout choices can move an agent score by several points.
- Benchmark quality. Reporting cites Epoch AI flagging SWE-bench Verified and DeepSWE v1.1 as problematic, which matters for a model whose RL environments include software-engineering tasks in the same family. Train-test overlap and environment similarity are legitimate questions when a lab trains on thousands of SWE-style tasks and then evaluates on SWE-style benchmarks. Nothing I found demonstrates contamination; the point is that the risk cannot be excluded from outside.
- Independent scores lag. Third-party numbers appeared within days, but sampling settings differ (the card recommends temperature 1.0, top-p 0.95).
For a broader treatment of how to read these suites, see our guide to AI agent benchmarks: SWE-bench, GAIA and tau-bench.
The 9B distill: a research checkpoint, with numbers
For the Qwen-based 9B, Xiaomi reports SWE-bench Verified 61.1 to 66.2, Cyber Bench 31.3 to 47.0, Terminal Bench 2.1 37.1 to 52.8, and Visual Coding 64.0 to 72.4 after its agentic RL. A 9B model gaining 15 points on Cyber Bench and Terminal Bench from RL on these environments is the strongest argument that the released environments have training value. It is also, again, a vendor-run result.
Access and Deployment: API Pricing and Self-Hosting Reality

Figure 4: Two ways in. Self-host the MIT weights with vLLM or SGLang, or call a hosted endpoint on Xiaomi’s platform, OpenRouter or DeepInfra at the listed tiers.
API pricing
Reported prices per million tokens, from the DataNorth summary and corroborated by VentureBeat and SiliconANGLE (via OpenRouter, 1.05M-token window):
| Tier | Input | Output |
|---|---|---|
| Flash | $0.14 | $0.28 |
| Pro | $0.435 | $0.87 |
| Pro-UltraSpeed | $4.35 | $8.70 |
DataNorth also lists Pro cached input at $0.004 per million. Xiaomi says Pro costs roughly one twentieth to one sixtieth of comparable overseas models, a claim tied to competitor list prices that I have not independently checked.
Worked cost example
Illustrative workload: an agent fleet consuming 100 million input tokens and 20 million output tokens per day, no caching.
- Flash: 100 times $0.14 plus 20 times $0.28 equals $14.00 plus $5.60, or $19.60 per day.
- Pro: 100 times $0.435 plus 20 times $0.87 equals $43.50 plus $17.40, or $60.90 per day.
- UltraSpeed: 100 times $4.35 plus 20 times $8.70 equals $435 plus $174, or $609 per day.
Agent workloads are input-heavy because every turn resends context, so the cached-input rate matters more than the headline. If 80% of input tokens hit the cache at $0.004, the Pro input line falls from $43.50 to roughly 20 times $0.435 plus 80 times $0.004, or $8.70 plus $0.32, about $9.02. Total is about $26.42 per day. Cache hit rates in practice depend on the host’s caching policy; this is a best-case illustration.
UltraSpeed is a latency product, not a cost product. At ten times the price for up to 20 times the speed, it is only rational where a human waits on the loop or a wall-clock deadline exists.
Self-hosting: what the model cards say
The model cards give launch commands and no hardware guidance; DataNorth notes explicitly that Xiaomi provides no self-hosting hardware sizing. What is published:
# Flash
vllm serve XiaomiMiMo/MiMo-V2.6-Flash-RL --tensor-parallel-size 4 --trust-remote-code
sglang serve --trust-remote-code --model-path XiaomiMiMo/MiMo-V2.6-Flash-RL --tp 8 --dp 2
# Pro
vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL --tensor-parallel-size 8 --trust-remote-code
sglang serve --trust-remote-code --model-path XiaomiMiMo/MiMo-V2.6-Pro-RL --tp 16 --dp 2 --ep 16 --nnodes 2
The Flash card lists tensor formats F32, BF16, F8_E4M3 and U8, so FP8 weights appear to be part of the release. The Pro card excerpt I read did not mention quantized variants.
Memory arithmetic (estimates)
Weight memory is parameters times bytes per parameter. These are my estimates, not Xiaomi figures, and they exclude KV cache, activations, and framework overhead.
- Pro, BF16: 1.02 trillion times 2 bytes is about 2.0 TB. FP8: about 1.0 TB.
- Flash, BF16: 309 billion times 2 is about 620 GB. FP8: about 310 GB.
A single 8-GPU node with 141 GB per GPU holds about 1.13 TB, so Pro in FP8 would barely fit weights with little headroom, which is consistent with the vLLM command using tensor parallelism of 8 and the SGLang command going to two nodes for expert parallelism. Flash at FP8 fits comfortably on 4 large GPUs at roughly 78 GB of weights each before cache. The commands’ parallelism (4 for Flash, 8 to 16 for Pro) matches the Winbuzzer report that Flash needs at least four-way and Pro eight- to sixteen-way parallel deployment.
Because only 10 of Pro’s 70 layers hold context-scaling KV cache, long-context serving is far more feasible than for an all-global trillion-parameter model, though 1M-token prefill remains expensive in time.
For the broader cost picture across self-hosting versus API, see AI inference cost optimization.
Licensing checklist
MIT permits commercial use, modification, and self-hosting with attribution. Practical items to verify before shipping: that the specific repository revision you pin is still MIT and ungated; that the 9B distill’s upstream Qwen license terms allow your use; and that the audio and vision encoders, which are part of the checkpoints, do not carry separate notices. For supply-chain hygiene on downloaded weights, see AI model supply chain security and provenance.
Limitations, Safety, and Failure Modes
Vendor-supplied evidence dominates. Most performance claims trace to Xiaomi’s own evaluation harnesses. The one independent anchor is the Artificial Analysis index, which measures a composite, not your workload. Independent replication of the RL-attribution claims does not exist yet.
Long-context degradation is a live risk. With 128-token windows in 86% of Pro’s layers (60 of 70), the model leans on 10 global layers for any dependency further than a few hundred tokens away. The 1M-token window is a capacity claim, not a quality guarantee. I found no published needle-in-haystack or multi-hop retrieval curve across the full window in the sources I could read. Test with your own documents at 100K, 500K, and 1M before relying on it.
Tool-call repetition. Xiaomi itself documents the failure where an agent keeps issuing similar tool calls without progress, and built MOPD2 to reduce it. The mitigation is applied to the Flash checkpoint that the card describes; whether every released checkpoint has the same protection is unclear. Wrap agents with loop detectors and step budgets regardless.
Reward hacking and grader dependence. Models trained against 7,000-odd environments with rubric judges and visual graders can learn to satisfy the grader. Ranking among successful trajectories improves signal but also concentrates optimization pressure on whatever the ranker prefers. Environments with executable tests are harder to game than rubric-judged ones, and the released dataset contains both kinds.
Cybersecurity capability cuts both ways. Xiaomi reports CyberGym 94.0 (Pro) and 95.1 (Flash), MiMo Cyber Bench 80.2, and SEC Bench Pro 66.3, and its RL mix deliberately includes 1,000 vulnerability-reproduction tasks. The same skill that helps defenders reproduce and triage bugs helps attackers. With MIT weights there is no server-side control, so any misuse mitigation is yours to build. I found no published red-team or safety-evaluation report in the material I reviewed; if one exists in the technical report, read it before deploying in a security-sensitive setting.
Jailbreak posture unknown. Open weights can be fine-tuned to remove refusals within hours. Do not assume the release-time behavior persists in a community derivative.
Serving complexity. Pro needs multi-GPU, often multi-node, deployment, and MoE serving is sensitive to expert load imbalance and all-to-all interconnect. The cards give commands but no throughput or latency numbers, so capacity planning requires your own load test.
Multimodal claims need scoping. Text, image, video and audio are all listed, yet the agent benchmarks are almost entirely text or GUI tasks. One outlet advises testing only the modalities you need, which is sound.
Reproduction gap. Even with environments and a verl fork, matching a run that used 25,088 rollouts per step at up to 1M context needs infrastructure few teams have. The realistic reproduction target is the 9B distill, not Pro.
How MiMo-V2.6 Compares
I am deliberately conservative here. I did not retrieve verified spec sheets for every peer, so the matrix is qualitative and grounded in what the sources report about MiMo, with pointers to our sibling deep-dives for peer specifics. Check each peer’s model card and license before deciding.
| Use case | MiMo-V2.6 Pro | MiMo-V2.6 Flash | Peer open models |
|---|---|---|---|
| Autonomous coding and terminal agents | Strong reported scores (Terminal Bench 2.1 at 89.9); heavier to serve | Near-Pro reported scores at about a third of the price | Compare on your harness: see DeepSeek V4 and Kimi K3 |
| Cheap high-volume API calls | $0.435/$0.87 per M tokens | Best value at $0.14/$0.28 | Check current pricing for GLM-5.2 and Qwen 3.6 |
| Self-hosting on modest hardware | Multi-node class | Four large GPUs class (estimate) | Smaller dense or MoE peers may fit a single node |
| Long-context repository work | 1M window, mostly-local attention; validate recall | Same window, smaller model | Test all candidates at your context length |
| Building your own agent RL | Study weights plus environments | Cheaper RL target | The 9B distill is the only small released checkpoint |
| Permissive licensing | MIT | MIT | Licenses differ across vendors; verify each |
The AI models pillar on this site groups these deep-dives; the DeepSeek V4 comparison is the closest architectural cousin in scale, while the Qwen 3.6 post covers the family whose 9B base Xiaomi used for its distill.
A note on the “MiMo vs DeepSeek V4” framing. Outlet comparisons put Pro ahead of DeepSeek V4.1 models on the Artificial Analysis index (46 against 36 for V4.1 Pro in one report). An index score compresses reasoning, coding, and agentic evaluations into one number, and it is a poor guide to any single workload. Run both on your own tasks.
Practical Recommendations
If you are choosing between Pro and Flash, start with Flash. The reported gap on agent benchmarks is small (DeepSWE 67.9 against 71.9, Terminal Bench 2.1 87.6 against 89.9), the price is about a third, and it runs on far less hardware. Move to Pro only where your own evaluation shows a measurable gain. Reserve UltraSpeed for interactive loops where latency has a dollar value.
If you self-host, budget for a serving stack before a model: interconnect, expert parallelism, observability, and a load test on your traffic shape. Pin the Hugging Face revision hash so a repository change cannot alter behavior or licensing under you.
If your interest is the RL stack, do not begin with Pro. Pull the RL environment dataset, run a subset against the 9B distill with the verl fork, and measure whether your gains match Xiaomi’s reported pattern. That is a weekend of exploration for a small team, not a six-day cluster commitment.
Checklist before adopting MiMo-V2.6:
- Pin exact model and dataset revisions; record license text at that hash.
- Run your own evaluation with fixed sampling (start from temperature 1.0, top-p 0.95 as the card suggests).
- Test recall at 100K, 500K, and 1M tokens on your documents.
- Add step budgets and repeated-call detection to every agent loop.
- Decide your misuse policy for cyber-capable weights before deployment.
- Compute cost with cached-input rates, not headline rates.
- Re-check third-party index rankings on the day you decide.
Frequently Asked Questions
Is MiMo-V2.6 really open source?
The Pro and Flash weights are open under the MIT license and downloadable from Hugging Face, which meets the common open-weight bar. The RL environments (Apache 2.0) and a verl fork are also public. Pretraining data and code are not part of the release, so it is open weight plus an open post-training toolkit rather than fully reproducible end to end.
How big is MiMo-V2.6 and what hardware does it need?
Pro has 1.02 trillion total parameters with 42 billion active; Flash has 309 billion with 15 billion active. Xiaomi publishes launch commands (tensor parallel 4 for Flash on vLLM, 8 to 16 for Pro, with two nodes on SGLang) but no formal hardware sizing. My estimate is about 1 TB of FP8 weights for Pro and about 310 GB for Flash, before cache.
How much does the MiMo-V2.6 API cost?
Reported per million tokens: Flash $0.14 input and $0.28 output, Pro $0.435 and $0.87, and Pro-UltraSpeed $4.35 and $8.70 for up to 20 times the output speed. Pro cached input is reported at $0.004. Prices are as relayed by OpenRouter-based coverage and Xiaomi says they are unchanged from V2.5; confirm on the provider before budgeting.
What did the MiMo-V2.6 reinforcement learning cost?
Reported RL compute was about $2.62 million for Pro and $850,000 for Flash, totaling $3.47 million, for 30 steps each in under six days. Each step used 1,568 prompts times 16 rollouts with contexts up to 1M tokens. Reporting says these figures exclude pretraining and other development work, so they are not the cost of building the model from scratch.
Is MiMo-V2.6 better than DeepSeek V4?
On the Artificial Analysis Intelligence Index, Pro scored 46 (46.32 on v4.3), which reporting places above DeepSeek V4.1 Pro at 36 in one comparison. That is one composite score from one time window. Your coding, retrieval, or tool-use workload may rank the models differently, so evaluate both on your own tasks before switching.
Can I use MiMo-V2.6 commercially?
The MIT license permits commercial use, modification, and redistribution with the license notice retained. Verify the license shown at the specific repository revision you download. The 9B distill is fine-tuned from Qwen3.5-9B, so also review the upstream Qwen license terms. This is not legal advice; have counsel confirm for regulated deployments.
Further Reading
- DeepSeek V4 architecture and benchmarks – the closest peer in scale.
- Kimi K3 explained – trillion-scale sparse MoE from Moonshot.
- GLM-5.2 explained – another open-weight agentic contender.
- Qwen 3.6 explained – the family behind the 9B distill base.
- Mixture-of-experts LLM architecture – MoE routing and load balance fundamentals.
- Xiaomi MiMo-V2.6 release notes – primary source from Xiaomi.
- MiMo-V2.6-Pro-RL on Hugging Face – model card with specifications and launch commands.
By Riju — about
