跳到正文
原文
OpenBMB· @OpenBMB · X·· 2 小时前AI 评分47
AI 导读

清华 NLP(OpenBMB 成员)联合中科院大学、东北大学、UIUC 和约翰霍普金斯大学提出 One-Shot OPD,将 OPD 训练集压缩到一条查询:数学任务上从 59.1 提升至 68.5(300 步),达到全量数据 OPD 69.8 的 87% 增益,并在代码、指令跟随和智能体工具使用上跨 Qwen、Llama、OLMo 成立。

正文

Post-training pipelines now use on-policy distillation (OPD) to hand a student the teacher's full next-token distribution at every prefix it visits: the student samples its own rollouts, and Qwen3, MiMo, GLM-5, DeepSeek-V4 and Kimi K3 all pair it with SFT and RL. Yet work on OPD has almost all stood on the algorithm side, treating the training set as given—so how much of OPD's gain does the data account for? Introducing One-Shot OPD, from @TsinghuaNLP (OpenBMB member) with the University of Chinese Academy of Sciences, Northeastern University, UIUC and Johns Hopkins University. It cuts the training set to one query, and the answer is that OPD is data-overfed but algorithm-starved.

1️⃣ One query, hundreds of steps. On math it goes from 59.1 to 68.5 by step 300, against 69.8 for full-data OPD—87% of its gain. It holds across code, instruction following and agentic tool use, and across Qwen, Llama and OLMo; a query the student never solves works about as well as one it always solves.
2️⃣ States, not questions. A prefix is a state where the teacher gives a target distribution, so 64 rollouts per step already yield tens of thousands of supervised positions. One query reaches 71.5% state coverage of what full-data OPD visits; 16 diverse queries reach 98.9% and match full data—extra queries buy new states, not new questions.
3️⃣Alignment slows, not the data. The student keeps improving, but each update absorbs less of what is left, which is why a run takes hundreds of steps rather than tens. This decline hardly depends on training-set size: on 1, 4, 16 and all 17k queries, alignment slowed at a similar pace. What limits a run is not how much data it gets, but how fast the student absorbs it.

📄 Paper: http://huggingface.co/papers/2609.04172
💻 Code: http://github.com/Thinking-Space/One-Shot-OPD
#AI #THUNLP #OpenBMB #LLM #PostTraining #Distillation #OpenSource

来源:OpenBMB · x.com