跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分46
AI 导读

Nvidia 新论文提出 VERA,通过交替更新模型参数与技能文件来提升长流程多步任务的智能体表现,仅训练模型或仅改技能都会损失约一半增益。VERA 构建超 9000 个可重启沙箱,用真实文件与日志逐步打分,决定下一步修复模型还是技能。在医学研究基准上,9B 智能体两种更新并用得分 69.1,仅改技能 56.1,仅训练 43.3。

正文

New Nvidia paper shows agents for long, multi-step work improve most when you alternate between training the model and editing its skill files, using step-by-step scores from real evidence to pick each fix.

Training only the model or only the harness leaves about half the gain on the table, compared with updating both in alternating rounds.

Most environments score only the final result, which hides which step broke. VERA builds over 9,000 restartable sandboxes that check each step against real files and logs, then improves the agent in rounds.

VERA turns benchmark runs into over 9,000 restartable sandboxes where each step of a long workflow gets its own checklist score from real evidence.

On a medical research benchmark, a 9B agent scored 69.1 with both kinds of updates, versus 56.1 with skill edits alone and 43.3 with training alone.

Score each step against real artifacts, and let those scores decide whether the next fix goes into the model or its skills.

来源:Rohan Paul · x.com