Xiaomi opens MiMo-V2.6 for agentic RL tests
Xiaomi released and open-sourced MiMo-V2.6, pairing open weights with reinforcement-learning resources for agent, coding, cyber and multimodal tasks.
Xiaomi has released and open-sourced the MiMo-V2.6 model series, turning its latest agent and multimodal work into a public challenge to larger open-model labs.
The MiMo announcement says the V2.6 series includes Pro and Flash models, with native multimodal capability, six days of live reinforcement-learning training and unchanged API prices from the V2.5 series. Xiaomi also linked the release to a Hugging Face collection for MiMo-V2.6, where the Pro RL, Flash RL and Distill-Qwen-9B entries are listed.

The strongest claim is not that MiMo-V2.6 has closed the entire frontier gap. Xiaomi says MiMo-V2.6-Pro reaches 46 on the Artificial Analysis composite intelligence index, above several named open systems on its chart, while still trailing Claude Fable 5.1 and GPT-6 Astra. That framing matters: the pitch is an open-weight model that narrows capability and cost gaps, not a clean win over the best closed systems.
The release is really about RL scale
Xiaomi says MiMo-V2.6-Pro and Flash each completed 30 reinforcement-learning steps in less than six days, with about 750,000 cumulative trajectories. The company reports training costs of about $2.62 million for Pro and $850,000 for Flash, along with task pass-rate improvements and a large jump on DeepSWE v1.1. Those are Xiaomi-reported figures, so they need independent replication before they become settled performance facts.
Still, the technical direction is important. Xiaomi describes larger batch sizes, a fully asynchronous architecture, one-million-token context support during training, more complex task environments across code, general, visual and cyber work, and a broader grader setup for long-horizon reinforcement learning. In plainer terms, it is trying to make agent training less like a one-off benchmark sprint and more like a repeatable production loop.

The benchmark table is useful because it also shows the unevenness. MiMo-V2.6-Pro looks strong on several general-agent and cyber rows, but it does not dominate every coding row, and some external frontier models remain ahead on selected tasks. That is the right way to read the release: not as a universal leaderboard crown, but as evidence that open agent models are becoming more specialized and more expensive to train well.
Open weights make the claim testable
The open-source section is the most consequential part for developers. Xiaomi says it has released weights, a technical report, a distill model, RL task environments and training framework resources built around tools such as verl, uni-agent and mini-swe-agent. The Hugging Face collection lists a 1T Pro RL model, a 311B Flash RL model and a 9B distill model.
That gives outside teams something to inspect beyond a launch chart. Researchers can test whether the model's reported agent gains hold across their own tasks, whether the training environments are reusable, and whether the smaller distill path is practical for labs that cannot run a trillion-parameter model. It also gives enterprise buyers a clearer comparison point when deciding whether an open-weight model is good enough for private coding, cyber or office workflows.
The caution is that open weights do not remove deployment work. MiMo-V2.6's largest models still need serious inference infrastructure, safety evaluation and task-specific routing. The interesting news is that Xiaomi is putting more of the agent-training machinery into public view. If the community can reproduce even part of the reported gains, open-model competition will move further from chat demos and deeper into full workflow systems.
Comments ()