跳到正文
Arena.ai· @arena · X·· 2 小时前AI 评分53
AI 导读

Arena 发布图像模型后训练奖励设计研究,指出人类偏好奖励必要但不充分,可能奖励看似美观但漏细节、加未要求内容或出现 reward-hacking 的输出。

正文

How to design rewards for post-training frontier image models?

Our research suggests human preference reward is necessary, but insufficient: A preference model may still reward outputs that look appealing but miss details, introduce unrequested content, or exhibit other forms of reward-hacking.

We therefore optimize towards a composite reward:
- Bradley-Terry reward model trained on ~5.6M pairwise human votes
- Faithfulness reward from auto-generated prompt checklists evaluated by a vision-language model
- Constraint reward covering explicit and implicit user intent
- Anti-reward-hacking rubric rewards targeting failures such as garbled text and photorealism drift
This post-training recipe improves two already-strong open image models:
- Post-trained FLUX.2-dev gains 69 Elo points on our live T2I leaderboard, scoring 1202
- Post-trained Ideogram 4 gains 20 Elo points reaching a score of 1224 and surpassing all publicly listed open models (as of Sep 04, 2026).

Offline ablations with Gemini 3.5 Flash as the judge, show that these reward components are complementary: win rate against the base model increases as we add faithfulness and then constraint rewards on top of preference-only training, reaching 64.2%.

Finally, we ensemble policies trained with and without the anti-reward-hacking objective directly in weight space, further increasing win rate to 66.0%.

来源:Arena.ai · x.com