面向拆分式 AI 评估的 Prediction-Powered Smoothing 与验证方法

HuggingFace Daily Papers(社区热门论文)·2026-09-17 08:00·5天前
AI 导读

研究提出 prediction-powered smoothing(PP-S)及跨报告分类体系借力的扩展 PP-TS,用于在标注样本有限时更准确地估计各领域评估均值。该方法还给出近似无偏的基于设计的交叉验证分数,用于在直接估计与平滑估计之间选择。在可验证评分的基准和人工标注的已部署智能体流量上,新估计器在点估计与区间估计上均优于直接估计器,覆盖率接近名义水平。

HuggingFace Daily Papers(社区热门论文)
40AI 编辑部评分,满分 100

面向拆分式 AI 评估的 Prediction-Powered Smoothing 与验证方法

2026-09-17 08:00· 5天前
AI 导读

研究提出 prediction-powered smoothing(PP-S)及跨报告分类体系借力的扩展 PP-TS,用于在标注样本有限时更准确地估计各领域评估均值。该方法还给出近似无偏的基于设计的交叉验证分数,用于在直接估计与平滑估计之间选择。在可验证评分的基准和人工标注的已部署智能体流量上,新估计器在点估计与区间估计上均优于直接估计器,覆盖率接近名义水平。

Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation.

For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, with an extension that borrows strength across a reporting taxonomy (PP-TS). For validation, we derive a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators. We study a curated benchmark with verifiable grading and deployed agent traffic graded by humans, each with every outcome observed. In both, the proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage.

At the same sampling budget, our score selects as well as an independent validation sample does and estimates the selected estimator's error far more accurately.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org