PRIMESCIENTIST shows that research agents should choose where to spend their experiments, rather than keep pushing the same idea until the budget runs out.
Instead of following a single trajectory, it keeps competing executable plans in a tree. Experiment results update branch values, while the allocation policy explores more when resources are plentiful and concentrates on stronger branches as the budget shrinks.
On 12 FIRE-Bench tasks, it delivered 10.3% higher average reward than AutoResearch while using 50.6% fewer research attempts under the same token budget. Across the full evaluation, it used fewer attempts on 23 of 24 tasks.