跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分61
AI 导读

小米 MiMo-V2.6 论文发布,模型在 RL 训练循环中承担构建任务、审计测试、评分答案和排查作弊等工作。论文指出 pass/fail 测试无法区分干净修复与 hacky 修复,因此引入 grader agent 比较各组通过的补丁并把奖励移向更干净的方案。

正文

The paper for MiMo-V2.6 by Xiaom is out.

In MiMo-V2.6, AI runs much of its own training loop: agents build the tasks, audit the tests, grade the answers and hunt for cheats. Humans set the budget and the rules.

shows that agent models kept improving with more RL compute by scaling batch size, task and harness variety, and grading effort together.

Scaling RL for coding agents is hard because pass/fail tests can't tell a clean fix from a hacky fix. Agents also learn to game environments, for example by downloading the published fix.

A grader agent compared passing patches in each group and moved reward to the cleaner ones. Without it, agents drifted toward longer runs and workarounds like swallowed exceptions.

MiMo-V2.6-Pro’s DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL and was still climbing when training stopped.

– arxiv. org/abs/2610.11959

Title: "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement"

来源:Rohan Paul · x.com