跳到正文
elvis· @omarsar0 · X·· 2 小时前AI 评分52
AI 导读

Microsoft 与加州大学圣巴巴拉分校等发布论文 ScholarEvolve,提出基于已发表智能体研究而非自身失败日志来进化 agent harness。方法将 harness 拆为工具使用、记忆管理和任务执行模块,对近期论文做主题建模以识别各模块的改进策略,再实现并测试组合,且可持续纳入新论文。

正文

New paper from Microsoft and colleagues on evolving agent harnesses.

It's a really cool idea to evolve a harness from published research. Something I have also been testing for the past couple of months.

ScholarEvolve proposes harness changes based on published agent research rather than the agent's own failure logs.

It splits the harness into modules for tool use, memory management, and task execution.

It runs topic modeling over recent papers to identify distinct improvement strategies for each module, then implements and tests combinations.

New papers can be added over time.

With the model held fixed, Qwen3.5-27B goal completion on AppWorld Challenge rises from 49.6% to 63.6%, and GPT-5.4-mini on Tau2-Bench Telecom rises from 72.7% to 81.9%.

Paper: https://arxiv.org/abs/2609.40169

Chat with Paper: https://academy.dair.ai/papers/learning-from-research-toward-lifelong-agent-harness-evolution-2609.40169

来源:elvis · x.com