We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agent to learn from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls (even when the overall trajectory was successful). The method combines rejection sampling fine‑tuning (RFT) with hint‑guided self‑distillation, and it cuts tool call failures by about 21% in live A/B tests
AI 导读
Perplexity 公布其 Computer 智能体后训练方法,通过模仿优质轨迹并显式纠正可避免错误(如错误工具调用),结合拒绝采样微调(RFT)与提示引导自蒸馏,在线上 A/B 测试中将工具调用失败率降低约 21%。引用推文补充,较晚训练的 checkpoint 相对早期 checkpoint 工具调用失败率下降 21.2%。
40
AI 编辑部评分,满分 100Perplexity 公布其 Computer 智能体后训练方法,通过模仿优质轨迹并显式纠正可避免错误(如错误工具调用),结合拒绝采样微调(RFT)与提示引导自蒸馏,在线上 A/B 测试中将工具调用失败率降低约 21%。引用推文补充,较晚训练的 checkpoint 相对早期 checkpoint 工具调用失败率下降 21.2%。
New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation. In a live A/B test, a later trained checkpoint ...
来源:Aravind Srinivas· x.com