微软新论文发现,编程智能体团队不设中央管理者时,规模越大得分越高、完成越快。在 ProgramBench 200 个任务中最难的 5 个上,无主管团队从 1 个扩展到 128 个智能体,每一步平均得分都在提升。该方案名为 Agensh,每个智能体自行认领子任务、构建测试并合并到共享 Git 仓库,所有运行使用同一模型,论文未报告大团队的成本。
Turns out coding agents don't need a boss:
New Microsoft paper finds that bigger teams of coding agents score higher and get there sooner when agents claim their own tasks without a central manager.
Even on the 5 hardest of ProgramBench's 200 tasks, scaling a manager-free team from 1 to 128 agents raised the average score at every step.
so for big jobs, add agents and let them coordinate through shared tools with no lead agent.
Popular multi-agent coding tools send all work through 1 lead agent, which can only manage so many helpers.
Agensh drops the lead agent. Each agent claims a sub-task, builds and tests it, merges it into a shared Git repo, and logs findings on a shared board.
Every run used 1 model, and the paper does not report what large teams cost.
来源:Rohan Paul · x.com