跳到正文
DAIR.AI· @dair_ai · X·· 2 小时前AI 评分40
AI 导读

Meta 提出 RankEvolve,通过编译协议约束各研究阶段与门控,让 Claude Code 和 Codex 作为独立节点互相审查修复改动。在同等预算下,双产品组合将执行准确率从单一最佳产品的 45.8% 提升至 62.5%。在开源 HSTU 推荐模型上迭代 12 次,MovieLens-20M 的 NDCG@10 较已发表结果提升 4.48%。

正文

Must-read paper from Meta on reliable auto-research agents.

If you let coding agents run ML experiments, one silent bug like leaked eval data or a disconnected gradient can invalidate hours of training and every iteration built on it.

RankEvolve enforces each research phase and gate through a compiled protocol. It runs Claude Code and Codex as separate nodes that review and repair each other's changes.

At a matched budget, combining the two products raises execution accuracy from 45.8% for the best single product to 62.5%.

Over twelve iterations on the open-source HSTU recommender, it improved NDCG@10 on MovieLens-20M by 4.48% over the published result.

Paper: https://academy.dair.ai/papers/rankevolve-a-reliable-multi-agent-auto-research-harness-for-evolving-ranking-mod-2609.39551

来源:DAIR.AI · x.com