NYU与Amazon新论文提出SGUID方法,只保留训练早期和后期都持续提供有效信号的技能,蒸馏少量精选技能即可匹配甚至超越大11倍的技能库。在3个Qwen模型上,按主题匹配选出的技能中不到25%能提供有效信号;仅用6个精选技能,4个模型中有3个在数学竞赛测试上匹配或超越了30至71个技能的完整库。第二轮加入3个新技能后,Qwen3-8B从64.3%提升至66.3%。
Good paper for selecting your skill files.
Before distilling a skill bank, log which skills give a steady training signal and drop the rest.
New NYU and Amazon paper finds that distilling a few skills that keep producing a useful training signal matches or beats distilling a skill bank up to 11× larger.
Skills are short written tips, like a rule for counting cases, that a model absorbs by learning from a copy of itself that reads them. Picked by topic match, under 25% of them gave any useful signal across 3 Qwen models.
SGUID keeps only skills that help early in training and still help late. With 6 such skills, 3 of 4 models matched or beat the full bank of 30 to 71 skills on math contest tests. A 2nd round with 3 new skills lifted Qwen3-8B from 64.3% to 66.3%.
来源:Rohan Paul · x.com