简而言之:《经济学人》用世界价值观调查对 25 个前沿 AI 模型进行了评分。相较于训练与对齐选择,模型所属实验室是更弱的预测因素:Gemini 与 Qwen 是邻居,GPT-4o 与 DeepSeek R1 近乎双胞胎,而 DeepSeek R1 与 DeepSeek V4 Flash 则形同陌路。世界观在代码生成中不可见。在商业分析、预测、招聘与政策工作中,它是一个实时输入。
《经济学人》让 25 个前沿 AI 模型完成了世界价值观调查1,这份问卷自 1981 年以来一直在描绘 100 个国家的道德信念。对于这个 2x2 矩阵,有两个轴:第一,从传统(宗教)到世俗。第二,从以集体基本需求为核心的生存,到自我表达与个人主义。
大多数模型位于该地图的自我表达一半,考虑到训练数据,这合情合理。
令人意外的是,这些模型彼此相距甚远。Gemini 3.1 Flash Lite 与 Qwen 3.6 Flash 互为邻居,在自我表达上最为靠前。
GPT-4o 与 DeepSeek R1 近乎双胞胎,一个在旧金山训练,一个在杭州训练。
DeepSeek R1 与 DeepSeek V4 Flash 来自同一实验室,却位于世俗/传统轴的两端。
共享的训练数据和相似的标注者解释了这些近乎孪生的模型。不同的后训练选择解释了那些迥异的模型。Common Crawl 中 46% 是英文2,因此模型所模仿的基础声音是一个受过大学教育的美国网民。Anthropic 随后将 Claude 对齐到联合国人权宣言中的原则3,这从本质上就是一份自由主义文件。
Grok 独树一帜,是一个传统的独立派。
这种差异改变了采购清单。如今每一份企业模型的 RFP 都会对价格、延迟、上下文窗口和基准分数进行评分。世界观不在清单上。它应该被列入吗?
对于代码生成、SQL、日志解析和图像分类来说,这没问题。计算机程序没有政治立场。
一旦模型被用于特定市场的商业决策,它的世界观就是一个实时输入。营销文案、用户行为预测和客户支持语气都必须与目标人群的价值观相匹配。
AI 世界观从未被视为 AI 采购的一部分,但对于某些用例,它可能需要成为一个考量因素。
In short : The Economist scored 25 frontier AI models on the World Values Survey. Lab of origin is a weaker predictor than training & alignment choices : Gemini & Qwen are neighbors, GPT-4o & DeepSeek R1 are near-twins, & DeepSeek R1 & DeepSeek V4 Flash are strangers. Worldview is invisible in code generation. In business analysis, forecasts, hiring, & policy work, it is a live input.
The Economist ran 25 frontier AI models through the World Values Survey1, the questionnaire that has mapped the moral beliefs of 100 countries since 1981. For this 2x2, there are two axes : first, traditional (religious) to secular. Second, survival, with a focus on collective basic needs, to self-expression & individualism.
Most models sit in the self-expression half of the map, which makes sense given the training data.
Surprisingly, the models are far apart. Gemini 3.1 Flash Lite & Qwen 3.6 Flash sit as neighbors, furthest in self-expression.
GPT-4o & DeepSeek R1 are near-twins, one trained in San Francisco, one in Hangzhou.
DeepSeek R1 & DeepSeek V4 Flash come from the same lab but lie at opposite ends of the secular / traditional axis.
Shared training data & similar labelers explain the near-twins. Different post-training choices explain the strangers. Common Crawl is 46% English2, so the base voice a model imitates is a college-educated American online. Anthropic then aligns Claude to principles from the UN Declaration of Human Rights3, a liberal document by construction.
Grok is off on its own, a traditional independent.
This variance changes the shopping list. Every RFP for an enterprise model today scores price, latency, context window, & benchmark scores. Worldview is not on the list. Should it be?
For code generation, SQL, log parsing, & image classification, that is fine. A computer program has no politics.
The moment a model is used for business decisions in a specific market, its worldview is a live input. Marketing copy, predictions of user behavior, & customer support tone all have to match the values of the target demographic.
AI worldviews have never been considered as part of AI procurement, but for certain use cases, it may need to become a consideration.