Perplexity · @perplexity_ai · X·2026-09-23 04:21·1小时前
AI 导读

新研究:我们通过提示引导的自蒸馏,对一个 Computer 模型进行后训练,让它从自身错误中学习。 在一次线上 A/B 测试中,较晚训练的 checkpoint 相比早期 checkpoint 将工具调用失败率降低了 21.2%。

Perplexity@perplexity_ai
32AI 编辑部评分,满分 100
2026-09-23 04:21· 1小时前
AI 导读

新研究:我们通过提示引导的自蒸馏,对一个 Computer 模型进行后训练,让它从自身错误中学习。 在一次线上 A/B 测试中,较晚训练的 checkpoint 相比早期 checkpoint 将工具调用失败率降低了 21.2%。

New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation.

In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint.

来源:Perplexity· x.com