跳到正文
原文
OpenBMB· @OpenBMB · X·· 3 小时前AI 评分45
AI 导读

面壁智能联合清华 THUNLP 提出 Diffusion Reward Models(DRM),不再把人类偏好压缩成单一分数,而是学习完整奖励分布,保留分歧、不确定性与多种合理判断。

正文

A reward of “3” can mean two completely different things.
Everyone thinks a response is mediocre — or half the people love it while the other half hate it.
Most Reward Models cannot tell the difference.
Introducing Diffusion Reward Models (DRM): instead of collapsing human preference into a single score, DRM learns the full reward distribution, preserving disagreement, uncertainty, and multiple plausible judgments.
✨ Paper:https://arxiv.org/abs/2609.33803
🤗 Models: https://huggingface.co/Teburile/DRM
💻 GitHub: https://github.com/thunlp/DRM

Why it matters:
Human disagreement is structured, not just noise. On datasets with repeated annotations, judgments often form separated or polarized patterns. More importantly, as human disagreement increases, DRM’s learned reward distribution becomes increasingly multimodal.

The distribution is useful, not just descriptive. DRM can use distributional uncertainty to identify unstable reward decisions, and distribution-aware ranking improves Best-of-N selection beyond simply taking the mean reward.

Reward Models get their own test-time scaling. Instead of only spending more compute on generating more responses, DRM can keep the response fixed and sample its reward distribution more times. More reward samples give a more reliable estimate — a new scaling axis unavailable to deterministic scalar RMs.

And the gains survive RLHF. When used as the training-time reward, DRM improves downstream policy performance over scalar reward baselines, showing that the benefit is not limited to offline RM benchmarks.

The takeaway: Reward modeling may lose something important when it compresses every human judgment into one number.
“Everyone thinks this is average” and “people strongly disagree about this” should not look identical to a Reward Model.
DRM makes that difference visible — and usable.

来源:OpenBMB · x.com