Xiaomi’s MiMo-V2.6 Tackles Agents, Cyber And 3D Worlds In Fully Open-Source Release

Xiaomi has launched and open-sourced the MiMo-V2.6 sequence, its newest era of natively omnimodal AI models, positioning the discharge as a key step in scaling reinforcement studying on verifiable, complicated duties. The lineup contains two fashions — MiMo-V2.6-Pro, described as the corporate’s most succesful mannequin to this point, and the extra cost-efficient MiMo-V2.6-Flash — alongside a Pro-UltraSpeed variant providing output speeds as much as 20 instances sooner.
On the Artificial Analysis Intelligence Index (v4.3, September 2026), MiMo-V2.6-Pro scored 46.32, surpassing Kimi K3 and Qwen3.8 Max to turn into the highest-rated open-source mannequin. Notably, the brand new sequence retains the API pricing of its V2.5 predecessor, pushing the intelligence-versus-cost Pareto frontier outward at unchanged value.
Benchmark outcomes place the fashions competitively throughout disciplines. On DeepSWE v1.1, a long-horizon software program engineering take a look at, MiMo-V2.6-Pro scored 71.9, trailing DeepSeek V4.1 Flash (74.2), Claude Opus 5 and GPT 6 Astra (74.0 every), whereas MiMo-V2.6-Flash reached 67.9 — a dramatic enchancment over MiMo-V2.5-Pro’s 19.0. The sequence additionally posted sturdy outcomes basically agentic workflows, main Automation Bench v1.0.6 amongst frontier fashions besides DeepSeek V4.1 Flash, and scoring 94.0 and 95.1 on the cyber-focused CyberFitness center benchmark.
Reinforcement Learning at Scale, Streamed in Public
The technical basis of the discharge is a scaled RL coaching run that Xiaomi streamed dwell. In below six days, every mannequin accomplished 30 RL steps over roughly 750,000 trajectories, at prices of roughly $0.85 million (Flash) and $2.62 million (Pro). Average go charges on coaching duties rose by 25% and 12% in relative phrases, and held-out benchmark beneficial properties have been substantial: DeepSWE v1.1 scores improved by roughly 17 factors for Flash (48.8 to 65.68) and 14 factors for Pro (58.4 to 72.57), with RL proving sample-efficient and generalizing past the coaching distribution.
Scaling was pursued alongside three axes: bigger batches on a completely asynchronous structure (1,568 samples per replace, as much as 1 million context size, 3.5–3.7 billion tokens per step); a multi-task suite spanning coding, normal brokers, visible and cyber domains; and elevated grader compute utilizing relative comparisons for extra exact reward indicators. The staff additionally froze the router to suppress coaching drift and constructed a protection in opposition to reward hacking combining reward design, adversarial analysis, anomaly detection and cross-verification.
Beyond benchmarks, Xiaomi highlights utilized capabilities below what it calls “Vibe World” — extending natural-language programming from software program to interactive 3D worlds, recreation improvement, Blender-based 3D modeling, closed-loop robotic arm management, frontend and presentation design, and end-to-end video and music manufacturing. In analysis settings, MiMo-V2.6-Pro contributed to computational screening of novel MOF supplies for PFAS adsorption and assisted in formalizing the Li–Yorke theorem in Lean 4, producing over 6,000 traces of kernel-verified code with out Lean-specific post-training.
The fashions can be found by the MiMo API Platform, AI Studio, MiMo Desktop and OpenRouter, with pricing unchanged from V2.5. Xiaomi is open-sourcing the complete technical report, coaching environments and RL code for replica and additional analysis.
The submit Xiaomi’s MiMo-V2.6 Tackles Agents, Cyber And 3D Worlds In Fully Open-Source Release appeared first on Metaverse Post.

Two omnimodal fashions, advancing by scaled reinforcement studying