Nathan Lambert · @natolambert · X·2026-09-17 04:41·45分钟前
AI 导读

小米MiMo-V2.6正进行大规模强化学习训练,并公开了训练细节。其RL扩展了三个维度:计算(每步约2B tokens,1568 prompts × 16 rollouts,全异步)、环境与harness(多任务智能体RL,单次运行混合多种harness)以及评分计算(智能体组内信用分配,基于测试用例和评分标准的奖励)。相关细节将在未来几周逐步开源。

Nathan Lambert@natolambert
48AI 编辑部评分,满分 100
2026-09-17 04:41· 45分钟前
AI 导读

小米MiMo-V2.6正进行大规模强化学习训练,并公开了训练细节。其RL扩展了三个维度:计算(每步约2B tokens,1568 prompts × 16 rollouts,全异步)、环境与harness(多任务智能体RL,单次运行混合多种harness)以及评分计算(智能体组内信用分配,基于测试用例和评分标准的奖励)。相关细节将在未来几周逐步开源。

One of the coolest at-scale RL resources made public yet! You love to see it.

Fuli LuoNearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scale...

来源:Nathan Lambert· x.com