Thomas Wolf · @Thom_Wolf · X·2026-09-17 21:42·2小时前
AI 导读

小米团队称 MiMo-V2.6 目前正处于 RL 训练中途,探索 RL 的规模上限。其扩展了三个方向:算力(每步约 2B tokens,1568 prompts × 16 rollouts。

Thomas Wolf@Thom_Wolf
54AI 编辑部评分,满分 100
2026-09-17 21:42· 2小时前
AI 导读

小米团队称 MiMo-V2.6 目前正处于 RL 训练中途,探索 RL 的规模上限。其扩展了三个方向:算力(每步约 2B tokens,1568 prompts × 16 rollouts。

Impressive level of openness on such a large run

Fuli LuoNearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scale...