跳到正文
elvis· @omarsar0 · X·· 1 小时前AI 评分48
AI 导读

Google 提出 VeriHarness,将同一基座模型变为智能体验证器,专门处理 rollout 间的一致与分歧:分歧时查工作区证据裁决,一致时反向挑战找被遗漏的需求。在五个长时程基准上取得最优选择分数,配合证据修订在 Gemini 3.5 Flash 上提升 6.2 分、Claude Opus 4.8 上提升 6.4 分,并开源约 26,000 条 rollout。

正文

Banger paper from Google.

It's standard practice to sample several agent rollouts and trust the answers they agree on.

This Google paper shows that agreement can hide shared errors, while disagreement often points to the correct alternative.

VeriHarness turns the same base model into an agentic verifier with two jobs.

One resolves claims where rollouts disagree by checking workspace evidence.

The other challenges claims that every rollout agrees on and looks for requirements they all missed.

Across five long-horizon benchmarks, it gives the best selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.

The authors also release about 26,000 rollouts.

Paper: https://arxiv.org/abs/2610.00972

Chat with Paper: https://academy.dair.ai/papers/veriharness-scaling-agentic-verification-for-long-horizon-tasks-2610.00972

来源:elvis · x.com