跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分49
AI 导读

NVIDIA 新论文提出,与其给 AI 智能体更多选项,不如给它一个更好的裁判:让小型智能体每步生成 8 个候选命令,再由裁判模型挑选执行。用前沿强模型当裁判时,成功率从 50% 提升到 68%,且无需重新训练;若由小模型自评,提升则小得多。

正文

Nvidia's new paper tell before giving your agent more options, give it a better judge.

A single bad command can derail an AI agent, so have it draft a few options and let a smart judge pick before acting.

And in this way, You can make an AI agent far more reliable without retraining it.

AI agents that work in a terminal usually run the 1st command they come up with. A single bad move, like installing the wrong package, can throw off every step after it.

NVIDIA researchers had a small agent draft 8 options at each step and let a judge choose which to run. With a strong frontier model as the judge, its success rate jumped from 50% to 68%, no retraining needed.

When the small model judged its own drafts, the gains were much smaller. More options don't help much if the judge can't tell them apart.

来源:Rohan Paul · x.com