New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation.
In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint.
新研究:我们通过提示引导的自蒸馏,对一个 Computer 模型进行后训练,让它从自身错误中学习。 在一次线上 A/B 测试中,较晚训练的 checkpoint 相比早期 checkpoint 将工具调用失败率降低了 21.2%。
新研究:我们通过提示引导的自蒸馏,对一个 Computer 模型进行后训练,让它从自身错误中学习。 在一次线上 A/B 测试中,较晚训练的 checkpoint 相比早期 checkpoint 将工具调用失败率降低了 21.2%。
New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation.
In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint.
来源:Perplexity· x.com