面向事件预测的语言模型与先验胜任度门控池化

HuggingFace Daily Papers(社区热门论文)·2026-09-10 08:00·5天前
AI 导读

研究提出"胜任度门控"方法,按领域从已解决结果中估计语言模型的边际价值权重,向全局权重收缩并重新校准池化预测。在2,357个已解决二元问题和五个语言模型上,该方法将外部基线Brier从0.0771降至0.0732,显著优于全局组合,但在ForecastBench官方市场子集上无显著提升。

HuggingFace Daily Papers(社区热门论文)
39AI 编辑部评分,满分 100

面向事件预测的语言模型与先验胜任度门控池化

2026-09-10 08:00· 5天前
AI 导读

研究提出"胜任度门控"方法,按领域从已解决结果中估计语言模型的边际价值权重,向全局权重收缩并重新校准池化预测。在2,357个已解决二元问题和五个语言模型上,该方法将外部基线Brier从0.0771降至0.0732,显著优于全局组合,但在ForecastBench官方市场子集上无显著提升。

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights.

We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The gain remains significant under leakage controls and against a leakage-safe time-series prior on the pooled structured set, with separate evidence on FRED.

In contrast, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market. Across four Qwen models, verbal confidence does not reliably identify when the model outperforms the external forecast, while outcome-estimated competence supports better abstention decisions. These results provide a practical approach for selective model use based on measured marginal value.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org