NVIDIA 发布 Mid-Harness 论文,提出在模型与 harness 之间采样并验证候选动作再执行的测试时计算方法。TerminalBench-Lite 上,GPT-5.6 Sol 验证器从 8 个采样动作中选择时,Pass@1 从 50.0% 升至 68.0%;弱验证器下增加采样几乎无收益。
Banger paper from NVIDIA on test-time compute for terminal agents.
The finding is that you should sample several candidate shell commands, verify them before running one, and spend more on the verifier than on extra samples.
With a GPT-5.6 Sol verifier choosing among 8 sampled actions, TerminalBench-Lite Pass@1 rises from 50.0% to 68.0%. With a weak verifier, extra samples add almost nothing.
Mid-Harness leaves the generator and harness unchanged and works between them. When a small TMAX-9B model verifies its own candidates, pairwise comparison works best, and distilling the strong verifier into it helps further.
Combining action sampling with trajectory sampling reaches higher success at lower estimated token cost than sampling full trajectories alone.
来源:DAIR.AI · x.com