跳到正文
DAIR.AI· @dair_ai · X·· 2 小时前AI 评分43
AI 导读

微软提出 ActiveSaddler,通过自适应调整训练场景来优化 AI 智能体框架(harness)。该方法将反复出现的失败归类为失败模式,以非平稳赌博机方式追踪每种模式的学习价值,动态分配预算在复现已知弱点和探索新弱点之间。在相同优化器下,GAIA2 测试 Pass@1 提升 4.4 分,Terminal-Bench 2.0 提升 7.5 分。

正文

Great paper from Microsoft and colleagues on optimizing agent harnesses.

Current harness optimizers change how the harness is updated but keep the training scenarios fixed, so feedback keeps coming from tasks that stop being informative as the harness improves.

This work adapts the scenarios as well.

ActiveSaddler groups recurring failures into failure patterns and treats each pattern as an arm in a non-stationary bandit.

It tracks how much the harness is still learning from each pattern and splits the budget between revisiting known weaknesses and finding new ones.

With the same optimizer, test Pass@1 improves by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 compared with a fixed scenario order.

Paper: https://academy.dair.ai/papers/activesaddler-automated-curriculum-learning-for-agent-harness-optimization-2610.00906

来源:DAIR.AI · x.com