SettledResearch
Xiaomi's MiMo-V2.6 report details how it scaled RL to 1M-token contexts
Xiaomi's LLM-Core team posted the MiMo-V2.6 technical report on arXiv Oct. 8. It describes an omni-modal model family trained with asynchronous RL at about 1,568 samples and 2.7 to 3.7B tokens per step, up to 1M context, with groupwise agentic grading and a frozen MoE router. The abstract gives no benchmark scores and does not mention model weights.
- HEAT
- OUTLETS
- 0
- FIRST SEEN
- LAST SIGNAL
Primary source
arXivarxiv.org/abs/2610.11959Coverage · 1 article, oldest first
Wednesday, October 7, 2026