CMU 论文提出 harness learning,用 RL 训练一个提案模型读取任务、当前 harness 和执行报告后生成 harness 代码修改,以修改后 harness 的得分为奖励,solver 模型权重不变。
Banger paper from CMU on harness learning.
(bookmark it)
Also, pay attention to this important new AI engineering skill of improving agents by editing their harness code instead of their weights.
Seeing a huge shift towards this.
The authors train a proposer model with RL to read a task, the current harness and an execution report, then write a code edit to the harness.
The reward is the score of the revised harness. The solver model never changes.
A trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families it never saw in training. A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA.
Paper: https://arxiv.org/abs/2609.35738
Chat with Paper: https://academy.dair.ai/papers/harness-learning-enables-generalizable-test-time-adaptation-2609.35738
来源:elvis · x.com