elvis · @omarsar0 · X·2026-09-17 05:35·48分钟前
AI 导读

小米MiMo-V2.6正进行大规模RL训练,从三个维度扩展:计算(每步约2B tokens、1568 prompts×16 rollouts、全异步)、环境与harness(多任务智能体RL,单次运行混合多个harness)以及评分计算(组内智能体信用分配,结合测试用例与rubric奖励)。团队表示将在未来数周逐步开源细节,目前投入已超100万美元。

elvis@omarsar0
46AI 编辑部评分,满分 100
2026-09-17 05:35· 48分钟前
AI 导读

小米MiMo-V2.6正进行大规模RL训练,从三个维度扩展:计算(每步约2B tokens、1568 prompts×16 rollouts、全异步)、环境与harness(多任务智能体RL,单次运行混合多个harness)以及评分计算(组内智能体信用分配,结合测试用例与rubric奖励)。团队表示将在未来数周逐步开源细节,目前投入已超100万美元。

This should be the standard for building open-source AI.

$1M+ so far.

Fuli LuoNearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scale...

来源:elvis· x.com