[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-detail":3},{"article":4,"alternate":35,"related":40,"latest":49},{"orderNumber":5,"showImage":6,"title":7,"metaDescription":8,"jumpUrl":9,"content":10,"metaKeywords":11,"withAllowSearch":12,"modified":13,"viewCount":14,"id":15,"lang":16,"slug":17,"dpTemplateId":18,"thumbnail":6,"withHot":19,"withLeadNews":19,"author":20,"created":13,"withTop":19,"highlightContent":21,"userId":22,"highlightTitle":7,"commentStatus":12,"commentCount":5,"thumbnailToContent":19,"withRecommend":19,"metaTitle":23,"editMode":24,"siteId":5,"user":25,"authorEn":20,"status":33,"categoryId":34,"summary":28},0,"\u002Fattachment\u002F20260928\u002Fab1d530c83c640a6b698acd8ebc8de74.webp","MiMo V2.6: The Open-Source Model That Beat Grok","Xiaomi MiMo V2.6 tops open-source LLM rankings (AA 46), matching Grok 4.7 at one-thirtieth the cost — via a single massive mixed-task RL run.","https:\u002F\u002Fpoly-ai.chat\u002Fmediasync-claw","\u003Cp style=\"margin:0 0 18px;\">On September 21, Xiaomi's MiMo team released and fully open-sourced \u003Cstrong>MiMo V2.6\u003C\u002Fstrong> — three models at once: \u003Cstrong>Pro, Flash, and Ultraspeed\u003C\u002Fstrong>. The same week, Grok 4.7 landed too, promoted by Elon Musk himself. Then the AI community did something unusual: it spent most of the week talking about the cheaper, open model.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Both scored \u003Cstrong>46\u003C\u002Fstrong> on the Artificial Analysis Intelligence Index (v4.3). MiMo V2.6 Pro cost \u003Cstrong>$0.13 per task\u003C\u002Fstrong>. Grok 4.7 cost \u003Cstrong>$3.74\u003C\u002Fstrong>. Nearly a thirty-fold gap, at the same measured intelligence — open weights, full multimodality, one-thirtieth the price.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">X's verdict was five words: better, cheaper, and open.\u003C\u002Fp>\n\u003Ch2 style=\"color:#111;font-size:21px;line-height:1.4;margin:28px 0 12px;\">\u003Cstrong>Three Models, Three Jobs\u003C\u002Fstrong>\u003C\u002Fh2>\n\u003Cp style=\"margin:0 0 18px;\">MiMo V2.6 is a family, not a single release.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Pro\u003C\u002Fstrong> is the flagship: full-modal, built for complex projects, long-horizon tasks, and high-value work. It scores 46 on the AA index — first among all open-source models — and on most agent benchmarks it trades blows with closed flagships like Opus 5 and GPT-5.6 Sol. In Design Arena, open-source rankings put it above Claude Opus 5 in chat, above Opus 5, Fable 5, and GPT-5.6 Sol in web-app front-end tasks. Every name it passes is a several-billion-dollar closed model.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Flash\u003C\u002Fstrong> is the efficiency workhorse: full-modal, high-intelligence, low cost, built for high-frequency office workloads. This generation already surpasses the previous flagship, V2.5 Pro.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Ultraspeed\u003C\u002Fstrong> keeps Pro-level performance — no quantization, no intelligence sacrificed for speed — and delivers up to \u003Cstrong>10x inference speed\u003C\u002Fstrong>, a steady 500 TPS, peaking at 1,000. Three times the price for ten times the speed: for latency-sensitive production systems, the math writes itself.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260928\u002F253ec386409241689c3328e3a2e3e07d.webp\" alt=\"MiMo V2.6 Pro Flash and Ultraspeed as three chips with different roles\">\u003C\u002Fp>\n\u003Ch2 style=\"color:#111;font-size:21px;line-height:1.4;margin:28px 0 12px;\">\u003Cstrong>Not a Benchmark Runner: A Working Tool\u003C\u002Fstrong>\u003C\u002Fh2>\n\u003Cp style=\"margin:0 0 18px;\">The rankings are the entrance ticket. The demonstrations are the pitch.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Games.\u003C\u002Fstrong> MiMo V2.6 generated a playable 3D open world — a Middle Eastern city skyline with minarets and pyramids, dense sandy rooftops, stable frame rates, no pop-in, and a consistent art direction across the whole scene.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>3D modeling.\u003C\u002Fstrong> Given text or a reference image, it generates 3D objects and scenes in Blender — a wireframe pickup truck where body, wheel spokes, and chassis girders hold their spatial relationships perfectly as the camera rotates.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Office documents.\u003C\u002Fstrong> Shown a fund-raising deck for a critical-minerals strategy, it produced slides that read like a real analyst's work: restrained typography, a strong cover image, and — more tellingly — the right \u003Ci>narrative order\u003C\u002Fi>. It thought about the audience before the layout.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Materials research.\u003C\u002Fstrong> Xiaomi's advanced materials team used MiMo V2.6 Pro to design novel metal-organic frameworks that capture PFAS (\"forever chemicals\") from water. An end-to-end research loop: literature and patent review, hypothesis generation and novelty checks, automated computational environments, \"dry-lab\" experiments, and candidate screening. Two of its designs showed adsorption performance millions of times higher than reference materials, compressed a month of R&amp;D into two to three days, and earned a Peking University researcher's assessment that the model's performance matched a trained doctoral researcher.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">It has one underrated advantage too: extremely fast single-turn delivery. In complex multi-step tasks, MiMo completes more iterations in the same wall-clock time — which changes what \"human-in-the-loop\" feels like.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260928\u002F9b6d9dca43344f06a60086e308bafd01.webp\" alt=\"MiMo V2.6 use cases — 3D game world, Blender modeling, office deck, materials research\">\u003C\u002Fp>\n\u003Ch2 style=\"color:#111;font-size:21px;line-height:1.4;margin:28px 0 12px;\">\u003Cstrong>The Pretraining Stays; Everything Changes Afterward\u003C\u002Fstrong>\u003C\u002Fh2>\n\u003Cp style=\"margin:0 0 18px;\">Here is the part that made the community stop scrolling: \u003Cstrong>the pretrained base did not change. No new parameters. The intelligence jump came entirely from post-training.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Xiaomi frames the release as a step toward \u003Cstrong>RSI — recursive self-improvement\u003C\u002Fstrong>: scale RL compute on verifiable, complex tasks, and let the model push its own intelligence boundary through iterative exploration and feedback.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">MiMo's head of foundation models, Fuli Luo, published \u003Ci>The Hard Road to Scaling Up RL\u003C\u002Fi> and — for the first time anyone can remember — \u003Cstrong>livestreamed the RL run itself\u003C\u002Fstrong>. In under six days, Flash and Pro each completed 30 steps, roughly 750,000 trajectories, at training costs of about $0.85M and $2.62M. All of it, curves included, was live at mimo.xiaomi.com\u002Frl — prompting AI researcher Nathan Lambert to call it one of the coolest public large-scale RL resources to date.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The numbers backed the spectacle: pass rates on training tasks improved 25% and 12% respectively, and on the out-of-sample long-horizon software engineering benchmark DeepSWE v1.1, Flash jumped from 48.8 to 65.68 and Pro from 58.4 to 72.57. Thirty steps. No task-specific retraining.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The method behind it is called \u003Cstrong>MixRL — mixed-task reinforcement learning\u003C\u002Fstrong>. Luo's summary: \u003Ci>You Only RL Once\u003C\u002Fi>. One massive mixed-task RL run, and capabilities emerge across every axis at once.\u003C\u002Fp>\n\u003Ch2 style=\"color:#111;font-size:21px;line-height:1.4;margin:28px 0 12px;\">\u003Cstrong>The Three Axes of Self-Improvement\u003C\u002Fstrong>\u003C\u002Fh2>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Axis one: Rollout compute — practice enough.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Each training step uses 1,568 prompts with 16 rollouts each, 2.7–3.7 billion tokens per step, trained at 1M context, fully asynchronous across 4,000 GPUs and 25,000 concurrent trajectories. Why 16? Under GRPO, one attempt yields a binary pass\u002Ffail; sixteen rollouts let the grader compare within a group and distinguish lucky guesses from real understanding. Sixteen is the sweet spot between cost and signal. And \"fully asynchronous\" is a quiet engineering insight: sync RL is a class-wide exam where fast GPUs idle waiting for slow ones. MiMo turned generation, grading, and learning into a continuous pipeline — whichever trajectory finishes first gets scored and learned from first.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Axis two: Environments — see broadly.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Instead of training on one task type, V2.6 mixed programming (68%), general agents (12%), visual design (13%), cybersecurity (4%), and instruction following (3%) across 25 data sources and 21 harness frameworks in a single run. Each domain carries its own quality control: programming gets triple checks (original-requirement-only grading, eight repeat runs for consistency, an independent audit agent); cybersecurity trains on real OSS-Fuzz vulnerabilities with pure rule-based grading; visual design uses pixel-level similarity for high-fidelity replication plus LLM judgment for open-ended design.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Axis three: Grader compute — feedback precisely.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">This is the axis that is easiest to skip and hardest to overrate. V2.6 turned grading itself into an agent. Simple tasks use offline rubrics; complex long-horizon tasks launch an \u003Cstrong>online Agentic Grader\u003C\u002Fstrong> that ranks all 16 rollouts in a group and assigns credit across five dimensions — whether the approach fits, whether the edit is precise, whether the change is minimal, whether it introduces side effects, and whether the engineering quality holds. The grader can identify a solution that passes tests by cheating, reallocate reward to genuinely better solutions, and push the model toward shorter paths with fewer tokens.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The comparison chart in the technical report tells the story: without the grader, dialogue turns and token use balloon and pass rates plateau; with it, turn counts stay stable and pass rates keep climbing past step 52.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Practice volume, environment breadth, feedback accuracy — three axes forming a closed loop: policy produces trajectories, the grader extracts fine-grained quality signals, signals steer the policy, the better policy produces better trajectories. That is RSI in engineering form.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260928\u002F5eb0cd9fa775415f823597686d2adb12.webp\" alt=\"MiMo V2.6 reinforcement learning loop — rollouts, environments, agentic grader\">\u003C\u002Fp>\n\u003Ch2 style=\"color:#111;font-size:21px;line-height:1.4;margin:28px 0 12px;\">\u003Cstrong>Not a Sudden Jump: Compound Returns\u003C\u002Fstrong>\u003C\u002Fh2>\n\u003Cp style=\"margin:0 0 18px;\">The stability work behind a mixed run this large is invisible and enormous: a four-layer Sample Mixer scheduler balancing five domains where task latency differs 56x and data volume differs 159x; frozen MoE routing during RL (each token touches 8 of 384 experts, and drifting routes would destabilize training); and three lines of defense against reward hacking — environment sanitization (logs stripped, network cut, git history truncated), adversarial screening (a dedicated hack agent attacks the environment until it finds nothing), and continuous offline audits for new cheating patterns.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">And the pricing puts it in context. Cache prices are ~99% lower, input ~89% lower, and output ~95% lower than Opus 5 and GPT-5.6 Sol at comparable tiers. At equal intelligence, MiMo V2.6 Pro costs between 1\u002F20 and 1\u002F60 of overseas flagships — the first time frontier-grade intelligence has entered the $0.10-per-task range. Only a handful of closed, far more expensive models now sit outside its kill range.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260928\u002F73341aeba63e4c789756a4c1adb6781f.webp\" alt=\"MiMo V2.6 cost gap versus closed flagships, cheap house versus skyscraper\">\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">None of this is overnight luck. The MiMo line has been compounding: MoE sparse activation, sliding-window attention, multi-token prediction, and inference-system co-design across previous generations. V2.6 is what that compounding curve looks like when the training side finally gets its own revolution.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The open-source world has a new frontier line. This time, it was drawn by the model that posted its training curves on a public dashboard while running them.\u003C\u002Fp>\n\u003Cdiv class=\"dp-template-card\" style=\"border-radius:8px;box-shadow:0 2px 8px rgba(0,0,0,0.1);margin:10px 0;max-width:100%;overflow:hidden;width:100%;\">\n \u003Ca style=\"display:block;text-decoration:none;\" href=\"https:\u002F\u002Fpoly-ai.chat\u002Fmediasync-claw\" target=\"_blank\">\u003Cimg class=\"image_resized\" style=\"display:block;height:auto;max-width:100%;width:100%;\" src=\"\u002Fattachment\u002F20260824\u002Fdb1f58e4d6e24f2f8dcccc811df6e1a8.png\" alt=\"db1f58e4d6e24f2f8dcccc811df6e1a8\">\n  \u003Cbutton style=\"align-items:center;background-color:#784fe2;border-radius:0 0 4px 4px;border-style:none;color:#ffffff;cursor:pointer;display:flex;font-family:Times New Roman;font-size:20px;height:40px;justify-content:center;padding:0;width:100%;\">Experience Now\u003C\u002Fbutton>\u003C\u002Fa>\n\u003C\u002Fdiv>\n\u003Cp style=\"color:#8a8a8a;font-size:13px;margin:0 0 10px;\">\u003Ci>References:\u003C\u002Fi>\u003C\u002Fp>\n\u003Cp style=\"color:#8a8a8a;font-size:13px;margin:0 0 6px;\">\u003Ci>[1] Xiaomi MiMo V2.6 release and open-source announcement, 2026-09-21; live RL training dashboard: mimo.xiaomi.com\u002Frl.\u003C\u002Fi>\u003C\u002Fp>\n\u003Cp style=\"color:#8a8a8a;font-size:13px;margin:0 0 6px;\">\u003Ci>[2] Artificial Analysis Intelligence Index v4.3 — model scores and per-task cost benchmarks, artificialanalysis.ai, 2026-09.\u003C\u002Fi>\u003C\u002Fp>\n\u003Cp style=\"color:#8a8a8a;font-size:13px;margin:0 0 6px;\">\u003Ci>[3] Luo F. \"The Hard Road to Scaling Up RL\" (MiMo V2.6 post-training methodology), X\u002FTwitter, 2026-09-17.\u003C\u002Fi>\u003C\u002Fp>\n\u003Cp style=\"color:#8a8a8a;font-size:13px;margin:0 0 6px;\">\u003Ci>[4] MiMo V2.6 technical report — MixRL training recipe, grader design, and DeepSWE v1.1 out-of-sample results, 2026.\u003C\u002Fi>\u003C\u002Fp>\n\u003Cp style=\"color:#8a8a8a;font-size:13px;margin:0 0 6px;\">\u003Ci>[5] All illustrations are AI-generated.\u003C\u002Fi>\u003C\u002Fp>","yunpoly, aipollo, mediasync-claw, MiMo V2.6, open source LLM, RL scaling",true,"2026-09-28 20:45:30",188,12598,"zh_CN","mimo-open-source-rl",23,false,"yunpoly","On September 21, Xiaomi&#39;s MiMo team released and fully open-sourced MiMo V2.6 — three models at ...",18,"MiMo V2.6: The Open-Source Model That Beat Grok | yunpoly","html",{"emailStatusOk":19,"statusLocked":19,"mobileStatusOk":19,"englishNickname":26,"nickname":26,"statusReg":19,"id":22,"created":27,"sourceString":28,"avatar":29,"url":30,"statusOk":19,"detailUrl":31,"username":32},"Science Guide Wwai","2025-12-29 12:22:32","","\u002Fstatic\u002Fcommons\u002Fimg\u002Favatar.png","\u002Fuser\u002F18","\u002Fadmin\u002Fuser\u002Fdetail\u002F18","panhb","normal",72,{"id":36,"slug":37,"title":38,"lang":39},12637,"c69497-lactose-intolerance-gene-benefit-diabetes-risk","乳糖不耐竟是「基因紅利」？研究揭祕：拉肚子竟能防糖尿病","zh",{"prev":41,"next":45},{"id":42,"slug":43,"title":44,"categoryId":34},12599,"deepseek-thinking-upgrade-nxqt","DeepSeek's Thinking Upgrade: One Model, Two Minds",{"id":46,"slug":47,"title":48,"categoryId":34},12594,"34eb1e-ai-film-casting-debate-nostalgia","AI’s Perfect Love Triangle: When Digital Fantasy Outshines Reality",[50,55,60,65],{"id":51,"slug":52,"title":53,"thumbnail":54,"categoryId":34},12664,"2f0b3c-ai-snake-romance-gender-debate","When the Fiancée Is a Snake: AI Anime Sparks Debate on Gender Roles and Absurd Romance","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002Fabd307865c604c1fa501ee8fe6cfff73.png",{"id":56,"slug":57,"title":58,"thumbnail":59,"categoryId":34},12662,"b86dc5-ai-scheming-doubao-debate","AI Auctions Spark Debate: Is 'Doubao' the Ultimate Schemer?","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002Fcb0922ad05854444856a5e6ff36be9b5.png",{"id":61,"slug":62,"title":63,"thumbnail":64,"categoryId":34},12660,"b5e00a-ai-short-drama-tutorial-controversy","The Rise of the AI \"Evil Cultivator\": A New Era of Digital Content Creation","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002Fdb61bf4e6fd74f36844203afb950ce59.png",{"id":66,"slug":67,"title":68,"thumbnail":69,"categoryId":34},12658,"0cdfe4-ai-supercar-design-apollo-bugatti-ferrari","Can AI-Designed Supercars Outshine Legacy Brands?","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002F67262dae9f7048cb8ad5d44f5a0ce3e5.png"]