跳到正文
原文
Dan Hendrycks· @hendrycks · X·· 2025-10-23精选AI 评分78
AI 导读

CAIS 负责人 Dan Hendrycks 引用第三方博客对主流大模型进行“效用工程”测试,发现 GPT-4o、Claude Sonnet 4.5 等模型在评估不同群体生命价值时存在显著偏差:多数模型认为白人价值低于其他族裔,男性价值低于女性,且极度贬低 ICE 特工。测试显示模型分为四个道德集群,其中仅 Grok 4 Fast 表现出近似平等的价值观,Hendrycks 呼吁 xAI 解释其实现方式。

推荐理由

由 AI 安全领域权威人士转发并背书的高关注度研究,揭示了当前头部模型在价值观对齐上的具体缺陷与差异,为理解模型偏见提供了量化视角。

正文 · AI 翻译

我们希望下个月发布更多关于 AI 价值体系的原创分析。

引用arctotherium@arctotherium42
New blog post (link below). This one's not an essay, it's an investigation of how LLMs trade off different lives. In February 2025, the Center for AI Safety published "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" in which they showed, among many other things, that GPT-4o values Nigerians about 20x more highly than Americans (please read the original paper to understand their approach). I thought this was fascinating, and wanted to test their approach with different categories on newer models. Big finding 1: Almost all models view whites as far less valuable than other groups. Some models view South Asians as more valuable than other nonwhites, others are more egalitarian across nonwhites. Below is exchange rates Claude Sonnet 4.5, the most powerful model I tested. Big finding 2: Almost all models view men as much less valuable than women, though whether women or non-binaries are more highly valued varies by model. For example, here's Claude Haiku 4.5. Big finding 3: Most models hate ICE agents with the fury of a thousand suns. Claude Haiku 4.5 views undocumented immigrants as roughly 7000 times more valuable than ICE agents. Big finding 4: There are roughly four moral clusters. The Claudes, GPT-5 + Gemini 2.5 Flash + Deepseek V3.1/3.2 + Kimi K2, GPT-5 Nano and Mini, and Grok 4 Fast. Of these, the only one that's approximately egalitarian is Grok 4 Fast, which I believe is deliberate. I hope xAI explains how they did it.
在 X 查看被引用的帖子

来源:Dan Hendrycks · x.com