Skip to content
LiveNext 14:05:55
SettledResearch

Xiaomi's MiMo-V2.6 report details how it scaled RL to 1M-token contexts

Xiaomi's LLM-Core team posted the MiMo-V2.6 technical report on arXiv Oct. 8. It describes an omni-modal model family trained with asynchronous RL at about 1,568 samples and 2.7 to 3.7B tokens per step, up to 1M context, with groupwise agentic grading and a frozen MoE router. The abstract gives no benchmark scores and does not mention model weights.

HEAT
OUTLETS
0
FIRST SEEN
LAST SIGNAL

Primary source

arXivarxiv.org/abs/2610.11959

Coverage · 1 article, oldest first