Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…
PROOF-Gen:从优化数据到更好的知识蒸馏
AI 导读
PROOF-Gen提出用优化后的数据改进工具调用能力的知识蒸馏。在τ²-bench上,教师模型57%的试运行失败,其中三分之二是近失(大部分工具调用正确),而传统生成-过滤流程因失败不提供信号,每轮都会遗留相同的难题。该方法通过利用失败信号优化数据,提升蒸馏效果。
Apple Machine Learning Research(RSS)
51
AI 编辑部评分,满分 100PROOF-Gen:从优化数据到更好的知识蒸馏
PROOF-Gen提出用优化后的数据改进工具调用能力的知识蒸馏。在τ²-bench上,教师模型57%的试运行失败,其中三分之二是近失(大部分工具调用正确),而传统生成-过滤流程因失败不提供信号,每轮都会遗留相同的难题。该方法通过利用失败信号优化数据,提升蒸馏效果。
原文 · 保持原样,未翻译原文 · 未翻译
来源:Apple Machine Learning Research(RSS)· machinelearning.apple.com