Stop making your AI agents debate and vote.
New Stanford+Together AI paper shows teams that learn to check each other's work solve problems none of them got right alone.
Many agent setups just have models debate, then vote. That mostly picks an answer someone already had.
Here, 1 model reviewed past team chats and rewrote the team's playbook, like who double-checks whom and who plays devil's advocate. Learning this took just 15 practice problems.
On math and physics tests, the 3-model team averaged 66.7%, versus 48.8% for its best model alone. It even beat perfectly choosing among the models' own answers, so teamwork created right answers none of them had.
So skip the debate-and-vote script: give agents clear jobs like checker and challenger, and let past runs improve them.