NVIDIA 新论文提出,与其给 AI 智能体更多选项,不如给它一个更好的裁判:让小型智能体每步生成 8 个候选命令,再由裁判模型挑选执行。用前沿强模型当裁判时,成功率从 50% 提升到 68%,且无需重新训练;若由小模型自评,提升则小得多。
Nvidia's new paper tell before giving your agent more options, give it a better judge.
A single bad command can derail an AI agent, so have it draft a few options and let a smart judge pick before acting.
And in this way, You can make an AI agent far more reliable without retraining it.
AI agents that work in a terminal usually run the 1st command they come up with. A single bad move, like installing the wrong package, can throw off every step after it.
NVIDIA researchers had a small agent draft 8 options at each step and let a judge choose which to run. With a strong frontier model as the judge, its success rate jumped from 50% to 68%, no retraining needed.
When the small model judged its own drafts, the gains were much smaller. More options don't help much if the judge can't tell them apart.
来源:Rohan Paul · x.com