Sakana AI 提出多智能体自监督方法 MASS,探索递归自改进
Recursive Self-Improvement through Collective Intelligence
Sakana AI 发布多智能体自监督(MASS):一个共享模型通过虚拟子代理团队解决任务、搜索更优多智能体工作流,并用自身判断选出最佳方案,再将团队经验蒸馏回共享模型。
Recursive Self-Improvement through Multi-Agent Self-Supervision
October 12, 2026Video 1 As AI systems tackle increasingly open-ended research problems, human experts may struggle to assess results and guide further progress. The model itself may be the best available optimizer and evaluator. The challenge is obtaining useful self-supervision within that homogeneous loop, where a model risks reinforcing its own blind spots rather than correcting them.
We ask whether a model can learn from the collective intelligence of its own team.
Introducing Multi-Agent Self-Supervision (MASS): Recursive self-improvement through collective intelligence:
- Blog: https://pub.sakana.ai/mass/
- Paper: https://arxiv.org/abs/2610.12176
- Code: https://github.com/SakanaAI/mass
We introduce Multi-Agent Self-Supervision (MASS). One shared model solves tasks through a team of virtual subagents, searches for better multi-agent workflows, then uses its own judgments to select the best ones. Training on the best team’s executions distills the team’s collective experience back into the shared model.
We ran two cycles using a 27B open-weights model on synthetic open-ended research tasks. After the second cycle, score per output token reached 1.2-1.6x the base model’s level across four research benchmarks.
Three key results:
- MASS enables role generalization within an RSI loop: training only on task-solving data from the optimizee naturally improves the capabilities of both the optimizer and the evaluator.
- multi-agent trajectories are more efficient training data than single-agent trajectories for RSI.
- optimizer capacity, rather than evaluator capacity, can be the bottleneck in the speed of improvement.
We believe the community needs to investigate homogeneous RSI carefully to support AI safety. When a model is also its own evaluator, shared blind spots could allow mistakes to be accepted as progress and reinforced through training. For AI to improve beyond human expertise safely, we need to understand these failures and preserve effective ways to detect them and intervene.
This work was done by Hyunin Lee, Jinglue Xu, Jeffrey Seely, Yujin Tang, and our collaborators from UC Berkeley (Somayeh Sojoudi, Matei Zaharia, Donghyun Lee).
来源:Sakana AI:Blog · sakana.ai