跳到正文
DAIR.AI· @dair_ai · X·· 3 小时前AI 评分47
AI 导读

Meta Superintelligence Labs 提出用专门的控制器决定智能体下一步该做什么,在相同 worker 和相同预算下,GPT-5.5 的 ProgramBench 成绩从 63.7% 提升到 71.5%,Codex 为 58.0%。

正文

Super interesting paper from Meta Superintelligence Labs on controlling long agent runs.

They use the same workers and same budget, and ProgramBench goes from 63.7% to 71.5% with GPT-5.5 when a dedicated controller decides what work to run next.

Codex scores 58.0%.

Here is how it works:

The controller keeps a short summary of the run and leaves full worker outputs in memory. Each cycle it updates that summary, proposes next steps, estimates what each one is worth under the remaining budget, and sends the chosen work to workers with the earlier outputs they need.

On ProofBench, ARC-AGI-2 and LongCoT-mini it adds 3.6 to 4.2 points over direct control, averaged over three frontier models.

Paper: https://academy.dair.ai/papers/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning-2609.38147

来源:DAIR.AI · x.com