Meta 提出 RankEvolve,通过编译协议约束各研究阶段与门控,让 Claude Code 和 Codex 作为独立节点互相审查修复改动。在同等预算下,双产品组合将执行准确率从单一最佳产品的 45.8% 提升至 62.5%。在开源 HSTU 推荐模型上迭代 12 次,MovieLens-20M 的 NDCG@10 较已发表结果提升 4.48%。
Must-read paper from Meta on reliable auto-research agents.
If you let coding agents run ML experiments, one silent bug like leaked eval data or a disconnected gradient can invalidate hours of training and every iteration built on it.
RankEvolve enforces each research phase and gate through a compiled protocol. It runs Claude Code and Codex as separate nodes that review and repair each other's changes.
At a matched budget, combining the two products raises execution accuracy from 45.8% for the best single product to 62.5%.
Over twelve iterations on the open-source HSTU recommender, it improved NDCG@10 on MovieLens-20M by 4.48% over the published result.
来源:DAIR.AI · x.com