On September 21, Xiaomi's MiMo team released and fully open-sourced MiMo V2.6 — three models at once: Pro, Flash, and Ultraspeed. The same week, Grok 4.7 landed too, promoted by Elon Musk himself. Then the AI community did something unusual: it spent most of the week talking about the cheaper, open model.
Both scored 46 on the Artificial Analysis Intelligence Index (v4.3). MiMo V2.6 Pro cost $0.13 per task. Grok 4.7 cost $3.74. Nearly a thirty-fold gap, at the same measured intelligence — open weights, full multimodality, one-thirtieth the price.
X's verdict was five words: better, cheaper, and open.
Three Models, Three Jobs
MiMo V2.6 is a family, not a single release.
Pro is the flagship: full-modal, built for complex projects, long-horizon tasks, and high-value work. It scores 46 on the AA index — first among all open-source models — and on most agent benchmarks it trades blows with closed flagships like Opus 5 and GPT-5.6 Sol. In Design Arena, open-source rankings put it above Claude Opus 5 in chat, above Opus 5, Fable 5, and GPT-5.6 Sol in web-app front-end tasks. Every name it passes is a several-billion-dollar closed model.
Flash is the efficiency workhorse: full-modal, high-intelligence, low cost, built for high-frequency office workloads. This generation already surpasses the previous flagship, V2.5 Pro.
Ultraspeed keeps Pro-level performance — no quantization, no intelligence sacrificed for speed — and delivers up to 10x inference speed, a steady 500 TPS, peaking at 1,000. Three times the price for ten times the speed: for latency-sensitive production systems, the math writes itself.

Not a Benchmark Runner: A Working Tool
The rankings are the entrance ticket. The demonstrations are the pitch.
Games. MiMo V2.6 generated a playable 3D open world — a Middle Eastern city skyline with minarets and pyramids, dense sandy rooftops, stable frame rates, no pop-in, and a consistent art direction across the whole scene.
3D modeling. Given text or a reference image, it generates 3D objects and scenes in Blender — a wireframe pickup truck where body, wheel spokes, and chassis girders hold their spatial relationships perfectly as the camera rotates.
Office documents. Shown a fund-raising deck for a critical-minerals strategy, it produced slides that read like a real analyst's work: restrained typography, a strong cover image, and — more tellingly — the right narrative order. It thought about the audience before the layout.
Materials research. Xiaomi's advanced materials team used MiMo V2.6 Pro to design novel metal-organic frameworks that capture PFAS ("forever chemicals") from water. An end-to-end research loop: literature and patent review, hypothesis generation and novelty checks, automated computational environments, "dry-lab" experiments, and candidate screening. Two of its designs showed adsorption performance millions of times higher than reference materials, compressed a month of R&D into two to three days, and earned a Peking University researcher's assessment that the model's performance matched a trained doctoral researcher.
It has one underrated advantage too: extremely fast single-turn delivery. In complex multi-step tasks, MiMo completes more iterations in the same wall-clock time — which changes what "human-in-the-loop" feels like.

The Pretraining Stays; Everything Changes Afterward
Here is the part that made the community stop scrolling: the pretrained base did not change. No new parameters. The intelligence jump came entirely from post-training.
Xiaomi frames the release as a step toward RSI — recursive self-improvement: scale RL compute on verifiable, complex tasks, and let the model push its own intelligence boundary through iterative exploration and feedback.
MiMo's head of foundation models, Fuli Luo, published The Hard Road to Scaling Up RL and — for the first time anyone can remember — livestreamed the RL run itself. In under six days, Flash and Pro each completed 30 steps, roughly 750,000 trajectories, at training costs of about $0.85M and $2.62M. All of it, curves included, was live at mimo.xiaomi.com/rl — prompting AI researcher Nathan Lambert to call it one of the coolest public large-scale RL resources to date.
The numbers backed the spectacle: pass rates on training tasks improved 25% and 12% respectively, and on the out-of-sample long-horizon software engineering benchmark DeepSWE v1.1, Flash jumped from 48.8 to 65.68 and Pro from 58.4 to 72.57. Thirty steps. No task-specific retraining.
The method behind it is called MixRL — mixed-task reinforcement learning. Luo's summary: You Only RL Once. One massive mixed-task RL run, and capabilities emerge across every axis at once.

