PRIMESCIENTIST:让研究智能体按预算分配实验

Rohan Paul · @rohanpaul_ai · X·2026-09-27 08:11·31分钟前
AI 导读

PRIMESCIENTIST 提出研究智能体应自主决定实验资源投向,而非沿单一思路耗尽预算。它用树结构维护多个可执行方案,实验结果更新分支价值,预算充足时多探索、预算收紧时聚焦强分支。在 12 项 FIRE-Bench 任务上,其平均奖励比 AutoResearch 高 10.3%,相同 token 预算下实验尝试次数少 50.6%,24 项任务中有 23 项尝试更少。

Rohan Paul@rohanpaul_ai
42AI 编辑部评分,满分 100

PRIMESCIENTIST:让研究智能体按预算分配实验

2026-09-27 08:11· 31分钟前
AI 导读

PRIMESCIENTIST 提出研究智能体应自主决定实验资源投向,而非沿单一思路耗尽预算。它用树结构维护多个可执行方案,实验结果更新分支价值,预算充足时多探索、预算收紧时聚焦强分支。在 12 项 FIRE-Bench 任务上,其平均奖励比 AutoResearch 高 10.3%,相同 token 预算下实验尝试次数少 50.6%,24 项任务中有 23 项尝试更少。

PRIMESCIENTIST shows that research agents should choose where to spend their experiments, rather than keep pushing the same idea until the budget runs out.

Instead of following a single trajectory, it keeps competing executable plans in a tree. Experiment results update branch values, while the allocation policy explores more when resources are plentiful and concentrates on stronger branches as the budget shrinks.

On 12 FIRE-Bench tasks, it delivered 10.3% higher average reward than AutoResearch while using 50.6% fewer research attempts under the same token budget. Across the full evaluation, it used fewer attempts on 23 of 24 tasks.

来源:Rohan Paul· x.com