跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分46
AI 导读

Meta、Duke 与加州大学的新论文提出将 Agent harness 的自动搜索拆分为多个专门分支,每个分支保留自己更擅长的问题并记录有效经验,再用路由器按任务分配最优 harness,在全部 4 个测试设置中均优于 Meta-Harness。

正文

New paper from Meta, Duke, California Univ on self-Improving Agent's Harness Optimization

When an AI tunes your agent's harness, split the search into specialized branches and route each task to the best fit, which beat Meta-Harness in all 4 test settings.

Giving each tuning branch its own problems and its own notes on what worked produced harnesses with different strengths, and a router turned those strengths into higher scores.

A harness is the code around an LLM that controls its tools, retrieval and self-checks. Meta-Harness has an AI rewrite it in a loop, but every version is scored on the same problems, so search sticks to 1 path.

Each of 2 branches keeps the practice problems it solves better than the other and writes its own notes on what worked. In math, 1 branch learned to verify answers, while the other learned to build full derivations.

With Gemini 3 Flash on Olympiad math, accuracy rose from 46.0% with Meta-Harness to 62.0%.

If you auto-tune agents, keep several specialized harnesses and route between them rather than betting on 1 winner.

来源:Rohan Paul · x.com