跳到正文
原文
DAIR.AI· @dair_ai · X·· 2 小时前AI 评分39
AI 导读

Salesforce AI Research 提出 Critical-State RL,用于多轮工具调用的强化学习训练:当奖励依赖后续轮次时,其变化很大程度来自下游噪声,该方法用嵌套采样分离出当前动作带来的奖励变化,只对选定的关键调用做上下文赌博机更新。在 BFCL v4 缺失函数任务上,训练选定的那一轮可提升约 14 分,而训练其他候选轮则准确率持平或下降。

正文

Great paper from Salesforce AI Research on RL for multi-turn tool use.

The finding is that you want to train the one call where the action changes the outcome, instead of spreading reward across the whole trajectory.

(bookmark it)

When reward depends on later turns, much of its variation comes from what happens downstream.

Critical-State RL uses nested sampling to separate the reward variation caused by the current action from that noise, then trains only the selected call with contextual-bandit updates.

On BFCL v4 missing-function tasks, training the selected turn adds about 14 points. Training the other candidate turn leaves accuracy flat or lower.

Paper: https://academy.dair.ai/papers/critical-state-rl-diagnosing-trainable-states-for-multi-turn-tool-use-2609.24985

来源:DAIR.AI · x.com