• Latest
  • Trending
  • All
Answer card: Xiaomi retrained MiMo-V2.6 to stop repeating tool calls, API swapped on 25 September 2026, MIT weights on 27 September, same model names.

MiMo-V2.6-Pro was quietly retrained to stop looping on tool calls

28 September 2026
Answer card: Anthropic committed $11.6 billion over seven years to Akamai Cloud for CPU workloads only, with revenue from the second half of 2027.

Anthropic’s $11.6B Akamai deal buys CPUs, not GPUs

27 September 2026
Answer card: Google Suncatcher MVP satellite, four Trillium TPUs on about one kilowatt, launching on SpaceX Transporter-18, reported for 1 October 2026.

Google’s first Suncatcher satellite flies four TPUs on 1 kW

25 September 2026
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

23 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 28, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

MiMo-V2.6-Pro was quietly retrained to stop looping on tool calls

by stephane
28 September 2026
in Dev
0
Answer card: Xiaomi retrained MiMo-V2.6 to stop repeating tool calls, API swapped on 25 September 2026, MIT weights on 27 September, same model names.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

An agent that greps the same path 148 times in one turn doesn't crash. It just looks busy until it hits max_tokens and you get the bill. That's the bug Xiaomi spent last week fixing in MiMo-V2.6. On 27 September it published retrained "MOPD" checkpoints of MiMo-V2.6-Pro and MiMo-V2.6-Flash, still under MIT, and said the API models had already been swapped on 25 September. Same names: mimo-v2.6-pro and mimo-v2.6-flash. If you run agents on either one, your model changed underneath you mid-week, and we think that's the part worth talking about.

The short answer

Xiaomi released MiMo-V2.6-Pro (1.02T parameters, 42B active) and MiMo-V2.6-Flash (310B, 15B active) as MIT open weights on 21 September 2026. Users then hit a failure mode where the model repeats the same tool call inside one turn. Xiaomi retrained both with a short extra distillation stage it calls MOPD, rolled that into the API on 25 September without renaming anything, and published the new weights on Hugging Face on 27 September as MiMo-V2.6-Pro-MOPD and MiMo-V2.6-Flash-MOPD. It says broader benchmarks held steady. If you self-host the original -RL checkpoints, switch, and send sampling parameters explicitly.

1.02%worst repetition rate before the fix, Flash in OpenCode
25 SeptAPI models swapped, names unchanged
574 GBof files in the Pro MOPD repo, by our count
Answer card stating that Xiaomi retrained MiMo-V2.6 to stop repeating tool calls, that the API models got the fix on 25 September 2026 and the MIT weights followed on 27 September without a name change, and that the worst pre-fix repetition rate was 1.02 percent of calls for Flash in the OpenCode harness.
A retrain that shipped to the API two days before anyone could read about it.

What was actually going wrong

MiMo-V2.6 landed on 21 September with a lot of attention. Artificial Analysis put Pro at 46 on its Intelligence Index, tied with Grok 4.7 and ahead of every other open weights model it tracks, per VentureBeat. API pricing's aggressive too: $0.435 input and $0.87 output per million tokens for Pro, $0.14 and $0.28 for Flash. So people wired it into coding agents fast.

And within a day, a GitHub issue on one agent harness described the model locking into the same tool call inside a single response until it ran out of tokens. The report cites 148 identical grep calls and 446 repeated bash checks in single turns, on the Flash API and on self-hosted MiMo-V2.6-Flash-RL under vLLM. The harness sent no sampling parameters at all, so vLLM fell back to near-greedy decoding. The workaround in the thread was checkpoint defaults (temperature 1.0, top_p 0.95) plus a repetition penalty of 1.05.

Xiaomi's own post on 27 September takes it seriously. It defines the metric as exact within-turn repetition, (N minus U) divided by N, where N is the tool calls in a turn and U the unique ones. Averaged over lots of turns, the numbers look tiny. Flash hit 1.02% in OpenCode, Pro 0.54% in the same harness, and most other harnesses sat well under 0.3%. Honestly, that framing undersells it. An average of 1% is mostly zeros plus the occasional turn that burns your whole output budget, and that tail is what you pay for.

Horizontal bar chart of exact within-turn tool-call repetition rates Xiaomi measured before the retrain: Flash at 1.02 percent in OpenCode, Pro at 0.54 percent in OpenCode, Flash at 0.27 percent in Claude Code, Pro at 0.19 percent in MiMo Desktop and Pro at 0.07 percent in Codex. After the retrain, many cells read 0.001 percent or less.
Figures from Xiaomi's technical blog of 27 September. Only five of its nine harnesses are shown.

What MOPD changed, and what Xiaomi didn't publish

MOPD is multi-teacher on-policy distillation. The model card for the new checkpoints describes a second version of it: several domain teachers (some trained with RL on verifiable tasks, some with supervised fine-tuning on synthetic demos) score the student's own continuations, including single new turns generated from partway through a teacher's or a demo's history. The repetition fix is described as a short run with a specialised teacher folded into that pass. Architecture's untouched. Same 70 layers, same 384 routed experts with 8 active, same 1M token context, and the same audio and vision encoders.

After the retrain, Xiaomi's heatmaps show many harness and context-length cells at 0.001% or below, or with no repeats observed. It says performance on the wider benchmark suite "held steady". We'd like to see that table. The card doesn't list post-MOPD benchmark numbers side by side with the RL ones, so for now it's Xiaomi's word, and the launch figures you've seen quoted (71.9 on DeepSWE v1.1, 89.9 on Terminal Bench 2.1) were measured on the -RL weights, not these.

The weights themselves are big but not new in size. The Pro MOPD repo holds about 574 GB of files by our count from the Hub API, identical in layout to the RL release, with the bulk of the tensors stored packed as U8 and BF16 plus FP8 for the rest. Flash is about 178 GB. Xiaomi's serving recipe for Pro is SGLang across two nodes with tensor parallel 16, or vLLM with tensor parallel 8. That's datacentre kit. For comparison, the 744B Atria Dawn preview ships 1.5 TB in BF16, so Xiaomi's packing does a lot of work here.

What we'd do if we ran it

If you call the API, you're already on the new weights, and Xiaomi hasn't mentioned an older snapshot you could pin. That's the uncomfortable bit. A silent swap that fixes a bug is still a silent swap, and if your evals were tuned on the 21 September behaviour, rerun them. Xiaomi did reset remaining MiMo Desktop quota as an apology, which is a nice touch, but API users got nothing beyond the fix itself.

If you self-host, move from -RL to -MOPD, and don't trust your harness to send sampling settings. Here's Xiaomi's vLLM line for Pro with one change from the GitHub thread: --generation-config auto, so the checkpoint's own defaults apply when a client sends nothing.

bash
vllm serve XiaomiMiMo/MiMo-V2.6-Pro-MOPD --tensor-parallel-size 8 --trust-remote-code --generation-config auto --reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice

And put a cap on identical tool calls in your agent loop regardless of model. It's ten lines of code. I'd argue it should've been there before any of this; loops like these aren't unique to Xiaomi, they're just better documented here than usual. I might be wrong about how rare they'll be on the MOPD weights in the wild, since Xiaomi's numbers come from its own harness runs. We'll check back once third party reports on the new checkpoints pile up.

Sources

Xiaomi MiMo, Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6, 27 September 2026 (metric, per-harness rates, API update on 25 September, quota reset). Hugging Face, MiMo-V2.6-Pro-MOPD model card and MiMo-V2.6-Flash-MOPD (MOPD2 method, architecture, serving commands, MIT licence; file sizes from the Hub API). VentureBeat, Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model, 21 September 2026 (parameters, pricing, Intelligence Index, launch benchmarks). GitHub, oh-my-pi issue 12784, opened 22 September 2026 (the repeated calls, sampling workaround).

Frequently asked questions

Is MiMo-V2.6-Pro-MOPD a new model?

No. It's the same architecture and size as MiMo-V2.6-Pro-RL with a short extra training stage, multi-teacher on-policy distillation, aimed mainly at cutting repeated tool calls. The API name mimo-v2.6-pro didn't change.

When did the API switch to the retrained weights?

On 25 September 2026, per Xiaomi's technical blog. The open weights followed on Hugging Face on 27 September. Xiaomi hasn't mentioned any way to keep calling the earlier version through the API.

Is the licence still MIT?

Yes. Both MOPD repos are tagged MIT, like the original RL checkpoints, with no revenue threshold in the model card.

Did the fix cost benchmark performance?

Xiaomi says broader benchmark performance held steady, but it hasn't published a side by side table. The widely quoted launch scores were measured on the earlier RL weights.

Tags: AI agentsllmMiMonewsopen-weightsXiaomi
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card: Xiaomi retrained MiMo-V2.6 to stop repeating tool calls, API swapped on 25 September 2026, MIT weights on 27 September, same model names.

MiMo-V2.6-Pro was quietly retrained to stop looping on tool calls

28 September 2026
Answer card: Anthropic committed $11.6 billion over seven years to Akamai Cloud for CPU workloads only, with revenue from the second half of 2027.

Anthropic’s $11.6B Akamai deal buys CPUs, not GPUs

27 September 2026
Answer card: Google Suncatcher MVP satellite, four Trillium TPUs on about one kilowatt, launching on SpaceX Transporter-18, reported for 1 October 2026.

Google’s first Suncatcher satellite flies four TPUs on 1 kW

25 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.