An agent that greps the same path 148 times in one turn doesn't crash. It just looks busy until it hits max_tokens and you get the bill. That's the bug Xiaomi spent last week fixing in MiMo-V2.6. On 27 September it published retrained "MOPD" checkpoints of MiMo-V2.6-Pro and MiMo-V2.6-Flash, still under MIT, and said the API models had already been swapped on 25 September. Same names: mimo-v2.6-pro and mimo-v2.6-flash. If you run agents on either one, your model changed underneath you mid-week, and we think that's the part worth talking about.
The short answer
Xiaomi released MiMo-V2.6-Pro (1.02T parameters, 42B active) and MiMo-V2.6-Flash (310B, 15B active) as MIT open weights on 21 September 2026. Users then hit a failure mode where the model repeats the same tool call inside one turn. Xiaomi retrained both with a short extra distillation stage it calls MOPD, rolled that into the API on 25 September without renaming anything, and published the new weights on Hugging Face on 27 September as MiMo-V2.6-Pro-MOPD and MiMo-V2.6-Flash-MOPD. It says broader benchmarks held steady. If you self-host the original -RL checkpoints, switch, and send sampling parameters explicitly.
What was actually going wrong
MiMo-V2.6 landed on 21 September with a lot of attention. Artificial Analysis put Pro at 46 on its Intelligence Index, tied with Grok 4.7 and ahead of every other open weights model it tracks, per VentureBeat. API pricing's aggressive too: $0.435 input and $0.87 output per million tokens for Pro, $0.14 and $0.28 for Flash. So people wired it into coding agents fast.
And within a day, a GitHub issue on one agent harness described the model locking into the same tool call inside a single response until it ran out of tokens. The report cites 148 identical grep calls and 446 repeated bash checks in single turns, on the Flash API and on self-hosted MiMo-V2.6-Flash-RL under vLLM. The harness sent no sampling parameters at all, so vLLM fell back to near-greedy decoding. The workaround in the thread was checkpoint defaults (temperature 1.0, top_p 0.95) plus a repetition penalty of 1.05.
Xiaomi's own post on 27 September takes it seriously. It defines the metric as exact within-turn repetition, (N minus U) divided by N, where N is the tool calls in a turn and U the unique ones. Averaged over lots of turns, the numbers look tiny. Flash hit 1.02% in OpenCode, Pro 0.54% in the same harness, and most other harnesses sat well under 0.3%. Honestly, that framing undersells it. An average of 1% is mostly zeros plus the occasional turn that burns your whole output budget, and that tail is what you pay for.
What MOPD changed, and what Xiaomi didn't publish
MOPD is multi-teacher on-policy distillation. The model card for the new checkpoints describes a second version of it: several domain teachers (some trained with RL on verifiable tasks, some with supervised fine-tuning on synthetic demos) score the student's own continuations, including single new turns generated from partway through a teacher's or a demo's history. The repetition fix is described as a short run with a specialised teacher folded into that pass. Architecture's untouched. Same 70 layers, same 384 routed experts with 8 active, same 1M token context, and the same audio and vision encoders.
After the retrain, Xiaomi's heatmaps show many harness and context-length cells at 0.001% or below, or with no repeats observed. It says performance on the wider benchmark suite "held steady". We'd like to see that table. The card doesn't list post-MOPD benchmark numbers side by side with the RL ones, so for now it's Xiaomi's word, and the launch figures you've seen quoted (71.9 on DeepSWE v1.1, 89.9 on Terminal Bench 2.1) were measured on the -RL weights, not these.
The weights themselves are big but not new in size. The Pro MOPD repo holds about 574 GB of files by our count from the Hub API, identical in layout to the RL release, with the bulk of the tensors stored packed as U8 and BF16 plus FP8 for the rest. Flash is about 178 GB. Xiaomi's serving recipe for Pro is SGLang across two nodes with tensor parallel 16, or vLLM with tensor parallel 8. That's datacentre kit. For comparison, the 744B Atria Dawn preview ships 1.5 TB in BF16, so Xiaomi's packing does a lot of work here.
What we'd do if we ran it
If you call the API, you're already on the new weights, and Xiaomi hasn't mentioned an older snapshot you could pin. That's the uncomfortable bit. A silent swap that fixes a bug is still a silent swap, and if your evals were tuned on the 21 September behaviour, rerun them. Xiaomi did reset remaining MiMo Desktop quota as an apology, which is a nice touch, but API users got nothing beyond the fix itself.
If you self-host, move from -RL to -MOPD, and don't trust your harness to send sampling settings. Here's Xiaomi's vLLM line for Pro with one change from the GitHub thread: --generation-config auto, so the checkpoint's own defaults apply when a client sends nothing.
vllm serve XiaomiMiMo/MiMo-V2.6-Pro-MOPD --tensor-parallel-size 8 --trust-remote-code --generation-config auto --reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice
And put a cap on identical tool calls in your agent loop regardless of model. It's ten lines of code. I'd argue it should've been there before any of this; loops like these aren't unique to Xiaomi, they're just better documented here than usual. I might be wrong about how rare they'll be on the MOPD weights in the wild, since Xiaomi's numbers come from its own harness runs. We'll check back once third party reports on the new checkpoints pile up.
Sources
Xiaomi MiMo, Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6, 27 September 2026 (metric, per-harness rates, API update on 25 September, quota reset). Hugging Face, MiMo-V2.6-Pro-MOPD model card and MiMo-V2.6-Flash-MOPD (MOPD2 method, architecture, serving commands, MIT licence; file sizes from the Hub API). VentureBeat, Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model, 21 September 2026 (parameters, pricing, Intelligence Index, launch benchmarks). GitHub, oh-my-pi issue 12784, opened 22 September 2026 (the repeated calls, sampling workaround).
Frequently asked questions
Is MiMo-V2.6-Pro-MOPD a new model?
No. It's the same architecture and size as MiMo-V2.6-Pro-RL with a short extra training stage, multi-teacher on-policy distillation, aimed mainly at cutting repeated tool calls. The API name mimo-v2.6-pro didn't change.
When did the API switch to the retrained weights?
On 25 September 2026, per Xiaomi's technical blog. The open weights followed on Hugging Face on 27 September. Xiaomi hasn't mentioned any way to keep calling the earlier version through the API.
Is the licence still MIT?
Yes. Both MOPD repos are tagged MIT, like the original RL checkpoints, with no revenue threshold in the model card.
Did the fix cost benchmark performance?
Xiaomi says broader benchmark performance held steady, but it hasn't published a side by side table. The widely quoted launch scores were measured on the earlier RL weights.






















