AgentWorld 将 3 到 20 个不同角色的 LLM 智能体放入游戏沙盒,执行 50+ 轮的长周期任务,智能体无法看到彼此内部状态,只能通过消息和共享计划协调。
More agents don't mean higher performance.
There is a coordination bottleneck to consider. Not to mention the unnecessary costs.
So how many of a multi-agent team's actions actually help it finish the task?
In this AgentWorld paper, fewer than a third.
AgentWorld puts 3 to 20 LLM agents with different roles into a game sandbox for tasks that run 50+ rounds.
Agents can't see each other's internal state, so they have to coordinate through messages and shared plans.
Gemini 3 Flash has the highest task success at 52.0%. Coordination tasks are the hardest category, at 12% success, and common failures include communication breakdowns, role confusion, and lost shared plans.
Paper: https://arxiv.org/abs/2609.31590
Chat with Paper: https://academy.dair.ai/papers/agentworld-benchmarking-long-horizon-collaboration-of-multi-agent-llms-2609.31590
来源:elvis · x.com