LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.
JEV-as-a-Judge:置信时接受,不确定时升级
AI 导读
JEV-as-a-Judge 用仅做决策的评审模型做低成本初筛,在普通偏好与证据支撑的事实性任务上,与最强对比的 SOTA LLM 评审差距在 3 个百分点以内,费用仅为其 0.36%。当判断需要检查推导过程或抵御精心编写的错误答案时,差距会扩大,且 JEV 与对比模型的差距集中在低置信度决策上。一个冻结的级联流程接受高置信度判定、升级不确定判定,以更低成本保留了对比模型 99% 的准确率。
HuggingFace Daily Papers(社区热门论文)
44
AI 编辑部评分,满分 100JEV-as-a-Judge:置信时接受,不确定时升级
JEV-as-a-Judge 用仅做决策的评审模型做低成本初筛,在普通偏好与证据支撑的事实性任务上,与最强对比的 SOTA LLM 评审差距在 3 个百分点以内,费用仅为其 0.36%。当判断需要检查推导过程或抵御精心编写的错误答案时,差距会扩大,且 JEV 与对比模型的差距集中在低置信度决策上。一个冻结的级联流程接受高置信度判定、升级不确定判定,以更低成本保留了对比模型 99% 的准确率。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org