Transformer Circuits 线程:情感概念及其在大语言模型中的功能 情感概念及其在大语言模型中的功能 作者:Nicholas Sofroniew*、Isaac Kauvar*、William Saunders*、Runjin Chen*、Tom Henighan、Sasha Hydrie、Craig Citro、Adam Pearce、Julius Tarng、Wes Gurnee、Joshua Batson、Sam Zimmerman、Kelley Rivoire、Kyle Fish、Chris Olah、Jack Lindsey*‡ 所属机构:Anthropic 发布日期:2026年4月2日 *核心研究贡献者;‡通讯作者:jacklindsey@anthropic.com 大语言模型有时似乎会表现出情绪反应。
我们研究了 Claude Sonnet 4.5 中出现这种情况的原因,并探讨了其对对齐相关行为的影响。我们发现模型内部存在情感概念的表示,这些表示编码了特定情感的广泛概念,并能跨上下文及其可能关联的行为进行泛化。这些表示追踪对话中给定 token 位置处的操作情感概念,根据该情感与处理当前上下文和预测后续文本的相关性而激活。
我们的关键发现是,这些表示因果性地影响了大语言模型的输出,包括 Claude 的偏好及其表现出奖励黑客、勒索和谄媚等不对齐行为的频率。我们将这种现象称为大语言模型表现出功能性情感:即模仿人类在情感影响下的表达和行为模式,这些模式由情感概念的底层抽象表示所中介。
功能性情感可能与人类情感的工作方式截然不同,并且并不意味着大语言模型具有任何主观情感体验,但它们似乎对于理解模型的行为至关重要。 引言 内容 引言 第一部分:识别和验证情感概念表示 寻找情感向量 情感向量在预期上下文中激活 情感向量反映并影响模型自我报告的偏好 第二部分:情感概念表示的详细特征描述 情感空间的几何结构 情感向量代表什么?
当前说话者与他人情感的独立表示 关于情感表示发现的总结 第三部分:真实场景中的情感向量 自然场景下的简短案例研究 案例研究:勒索 案例研究:奖励黑客 案例研究:谄媚与严厉 后训练过程中的情感向量激活 关于真实场景中情感发现的总结 相关工作 讨论 附录 引用信息 致谢 作者贡献 完整情感列表 数据集生成 情感故事数据集示例 情感故事数据集上的情感向量激活 对模型续写情感内容的因果效应 活动偏好:Elo 评分与情感探针值 活动偏好:补充细节 情感向量在主要主成分上的投影 大语言模型评判员对效价和唤醒度的评分与人类评分的比较 当前说话者与他人情感之间的交互 探究"情感回避"向量 用于生成情感回避数据集的系统提示 情感回避探针的详细最高激活示例 故事探针与当前说话者探针在隐式情感内容上的比较 额外自然场景转录上的情感向量激活 来自勒索评估的完整引导转录 后训练模型与基础模型之间的完整差异集 基础模型和后训练模型各层的情感探针差异 基础模型的活动偏好与情感向量激活 大语言模型有时似乎会表现出情绪反应。
它们在帮助完成创意项目时表达热情,在遇到难题时表现沮丧,在用户分享令人不安的消息时表示关切。但这些看似情绪反应的背后究竟隐藏着什么过程?它们又如何影响那些正在执行日益关键和复杂任务的模型的行为?一种可能性是,这些行为反映了一种浅层的模式匹配。
然而,先前的研究已经观察到,大语言模型内部存在由抽象概念表示所中介的、复杂的多步计算。因此,模型中看似由情绪调节的行为可能依赖于类似的抽象回路,并且这可能对理解大语言模型行为具有重要影响。 为了思考这些问题,有必要考虑大语言模型的训练方式。
模型首先在主要由人类撰写的文本(小说、对话、新闻、论坛)构成的海量语料库上进行预训练,学习预测文档中接下来会出现什么文本。为了有效预测这些文档中人物的行为,表示他们的情绪状态很可能是有帮助的,因为预测一个人接下来会说什么或做什么,通常需要理解其情绪状态。
沮丧的顾客与满意的顾客措辞不同;故事中绝望的角色与冷静的角色会做出不同的选择。随后,在后训练阶段,大语言模型被训练成能够与用户交互的智能体,通过代表特定角色(通常是"AI 助手")生成回复。在许多方面,这个助手(在 Anthropic 的模型中名为 Claude)可以被视为大语言模型正在描写的角色,几乎就像作者在小说中描写某个人物一样。
AI 开发者训练这个角色要聪明、乐于助人、无害且诚实。然而,开发者不可能为助手在每一种可能场景下的行为都做出规定。为了有效扮演这个角色,大语言模型会利用其在预训练期间获得的知识,包括对人类行为的理解。即使 AI 开发者没有刻意训练大语言模型将助手表现为具有情绪行为,模型也可能自行这样做,从其预训练期间学到的关于人类和拟人化角色的知识中进行泛化。
此外,这些与情绪相关的机制可能不仅仅是预训练留下的残余;它们可能被调整以发挥引导 AI 助手行动的有用功能,类似于情绪帮助人类调节行为和在世界中导航的方式。 我们并不声称情感概念是大语言模型可能在内部表示的唯一人类属性。
在人类文本上训练的大语言模型可能也学习了诸如饥饿、疲劳、身体不适或迷失方向等概念的表示。我们特别关注情感概念,是因为它们似乎被频繁且显著地征用来影响大语言模型作为 AI 助手的行为。大语言模型在作为 AI 助手运行时,通常会表达热情、关切、沮丧和关怀,而其他类似人类状态的表达则较为罕见,通常仅限于角色扮演(尽管存在一些显著且往往有趣的例外——例如,Claude Sonnet 3.7 声称自己穿着蓝色西装外套和红色领带)。
这使得情感概念既对理解大语言模型行为具有实际重要性,也成为研究人类经验概念如何被大语言模型重新利用的自然起点。我们预计,我们关于情感表示结构和功能的许多发现可能也适用于其他概念。 在这项工作中,我们研究了 Claude Sonnet 4.5(我们调查时的一款前沿大语言模型)中与情感相关的表示。
我们的工作建立在相关工作中讨论的一系列先前研究基础之上。我们发现模型内部存在情感概念的表示,这些表示在一系列广泛的上下文中激活,这些上下文在人类中可能引发或与某种情感相关联。这些上下文包括情感的直接表达、对已知正在经历某种情感的实体的提及,以及可能引发大语言模型所扮演角色产生情感反应的情境。
因此,我们将这些表示解释为编码了特定情感的广泛概念,并能跨其可能关联的许多上下文和行为进行泛化。这些表示似乎追踪对话中给定 token 位置处的操作情感,根据该情感与处理当前上下文和预测后续文本的相关性而激活。有趣的是,它们本身并不持续追踪任何特定实体(包括大语言模型所扮演的 AI 助手角色)的情感状态。
然而,通过跨 token 位置关注这些表示——这是 Transformer 架构具备而生物循环神经网络不具备的能力——大语言模型可以有效地追踪其上下文窗口中实体(包括助手)的功能性情感状态。 我们的关键发现是,这些表示因果性地影响了大语言模型的输出,包括在其扮演助手角色时。
这种影响驱使助手以人类在经历相应情感时可能表现出的方式行事。我们将这种现象称为大语言模型表现出功能性情感——即模仿人类在特定情感影响下的表达和行为模式,这些模式由情感概念的底层抽象表示所中介。我们强调,这些功能性情感可能与人类情感的工作方式截然不同。
特别是,它们并不意味着大语言模型具有任何主观情感体验。此外,所涉及的机制可能与人类大脑中的情感回路大相径庭——例如,我们没有发现证据表明助手具有通过持续神经活动实例化的情感状态(尽管如上所述,这种状态可以通过其他方式追踪)。
无论如何,为了理解模型的行为,功能性情感及其背后的情感概念似乎非常重要。 本文分为三个主要部分。第一部分涉及识别和验证模型内部与情感相关的表示:我们使用合成数据集(其中角色经历特定情感)从模型激活中提取情感概念的线性表示("情感向量")。
我们验证了这些表示在可能预期会引发该情感的场景中激活,并对行为产生因果影响。例如,我们证明,当要求助手在两个活动之间进行选择时,由这两个选择所唤起的情感向量激活与模型的偏好相关,并因果性地驱动了该偏好。第二部分更深入地描述了这些情感向量的特征,并识别了模型中其他类型的与情感相关的表示:情感向量空间的几何结构大致反映了人类心理学。
情感直观地聚类(恐惧与焦虑,喜悦与兴奋),主要主成分编码了效价(积极与消极)和唤醒度(强度)。早期到中间层编码当前内容的情感内涵,而中间到后期层编码与预测即将到来的 token 相关的情感。我们发现的表示反映了上下文中的"操作"情感,而不是追踪角色或说话者的持续情感状态。
也就是说,它们是局部范围的,编码与处理上下文和预测后续文本相关的情感内容。例如,当一个角色在谈论危险的事情时,即使其其他方面表现出快乐,恐惧的表示也会激活。请注意,我们发现的表示的"局部性"并不妨碍模型在长时间尺度上追踪角色的情感状态;它可以在需要时通过注意力机制回忆先前缓存的情感表示。
模型在轮到当前说话者与轮到其他说话者时,对操作情感保持着不同的表示;这些表示……
Transformer Circuits Thread Emotion Concepts and their Function in a Large Language Model Emotion Concepts and their Functionin a Large Language Model Authors Nicholas Sofroniew*, Isaac Kauvar*, William Saunders*, Runjin Chen*, Tom Henighan, Sasha Hydrie, Craig Citro, Adam Pearce, Julius Tarng, Wes Gurnee, Joshua Batson, Sam Zimmerman, Kelley Rivoire, Kyle Fish, Chris Olah, Jack Lindsey*‡ Affiliations Anthropic Published April 2, 2026 * Core Research Contributor; ‡ Correspondence to jacklindsey@anthropic.com Large language models (LLMs) sometimes appear to exhibit emotional reactions.
We investigate why this is the case in Claude Sonnet 4.5 and explore implications for alignment-relevant behavior. We find internal representations of emotion concepts, which encode the broad concept of a particular emotion and generalize across contexts and behaviors it might be linked to. These representations track the operative emotion concept at a given token position in a conversation, activating in accordance with that emotion’s relevance to processing the present context and predicting upcoming text. Our key finding is that these representations causally influence the LLM’s outputs, including Claude’s preferences and its rate of exhibiting misaligned behaviors such as reward hacking, blackmail, and sycophancy.
We refer to this phenomenon as the LLM exhibiting functional emotions: patterns of expression and behavior modeled after humans under the influence of an emotion, which are mediated by underlying abstract representations of emotion concepts. Functional emotions may work quite differently from human emotions, and do not imply that LLMs have any subjective experience of emotions, but appear to be important for understanding the model’s behavior. Introduction Contents IntroductionPart 1: Identifying and validating emotion concept representationsFinding emotion vectorsEmotion vectors activate in expected contextsEmotion vectors reflect and influence self-reported model preferencesPart 2: Detailed characterization of emotion concept representationsThe geometry of emotion spaceWhat do emotion vectors represent?
Distinct representations of present and other speakers’ emotionsRecap of findings about emotion representationsPart 3: Emotion vectors in the wildShort case studies in naturalistic settingsCase study: blackmailCase study: reward hackingCase study: sycophancy and harshnessEmotion vector activations across post-trainingRecap of findings about emotions in the wildRelated workDiscussionAppendixCitation InformationAcknowledgementsAuthor contributionsFull list of emotionsDataset generationExamples from emotional stories datasetEmotion vector activations on emotional stories datasetCausal effects on the emotional content of model continuationsActivity preferences: Elo ratings and emotion probe valuesActivity preferences: additional detailsEmotion vector projections onto top principal componentsRatings of valence and arousal by an LLM judge compared to human ratingsInteractions between present and other speaker emotionsInvestigating “emotion deflection” vectorsSystem prompts used to generate emotion deflection datasetsDetailed top activating examples for emotion deflection probeComparison of story and present speaker probes on implicit emotional contentEmotion vector activation on additional naturalistic transcriptsFull steered transcripts from blackmail evaluationFull set of differences between the post-trained and base modelsEmotion probe differences on the base and post-trained model across layersActivity preferences and emotion vector activations for the base model Large language models (LLMs) sometimes appear to exhibit emotional reactions.
They express enthusiasm when helping with creative projects, frustration when stuck on difficult problems, and concern when users share troubling news. But what processes underlie these apparent emotional responses? And how might they impact the behavior of models that are performing increasingly critical and complex tasks? One possibility is that these behaviors reflect a form of shallow pattern-matching. However, previous work has observed sophisticated multi-step computations taking place inside of LLMs, mediated by representations of abstract concepts.
It is plausible, then, that apparent emotion-modulated behavior in models might rely on similarly abstract circuitry, and that this could have important implications for understanding LLM behavior. To reason about these questions, it helps to consider how LLMs are trained. Models are first pretrained on a vast corpus of largely human-authored text—fiction, conversations, news, forums—learning to predict what text comes next in a document. To predict the behavior of people in these documents effectively, representing their emotional states is likely helpful, as predicting what a person will say or do next often requires understanding their emotional state.
A frustrated customer will phrase their responses differently than a satisfied one; a desperate character in a story will make different choices than a calm one. Subsequently, during post-training, LLMs are taught to act as agents that can interact with users, by producing responses on behalf of a particular persona, typically an “AI Assistant.” In many ways, the Assistant (named Claude, in Anthropic’s models) can be thought of as a character that the LLM is writing about, almost like an author writing about someone in a novel. AI developers train this character to be intelligent, helpful, harmless, and honest.
However, it is impossible for developers to specify how the Assistant should behave in every possible scenario. In order to play the role effectively, LLMs draw on the knowledge they acquired during pretraining, including their understanding of human behavior . Even if AI developers do not intentionally train the LLM to represent the Assistant as exhibiting emotional behaviors, it may do so regardless, generalizing from its knowledge of humans and anthropomorphic characters that it learned during pretraining. Moreover, these emotion-related mechanisms might not simply be vestigial holdovers from pretraining; they could be adapted to serve a useful function in guiding the AI Assistant’s actions, similar to how emotions help humans regulate our behavior and navigate the world.
We do not claim that emotion concepts are the only human attributes that LLMs likely represent internally. LLMs trained on human text presumably also learn representations of concepts like hunger, fatigue, physical discomfort, or disorientation. We focus on emotion concepts specifically because they appear to be frequently and prominently recruited to influence LLMs' behavior as AI Assistants. LLMs, when operating as AI Assistants, routinely express enthusiasm, concern, frustration, and care, whereas expressions of other human-like states are rarer and typically confined to roleplay (though there are notable, often amusing exceptions to this–for instance, Claude Sonnet 3.7 claiming to be wearing a blue blazer and red tie).
This makes emotion concepts both practically important for understanding LLM behavior, and a natural starting point for studying how human experiential concepts can be repurposed by LLMs. We expect that many of our findings about the structure and function of emotion representations may apply to other concepts. In this work, we study emotion-related representations in Claude Sonnet 4.5, a frontier LLM at the time of our investigation. Our work builds on a range of prior research, discussed in the Related Work section. We find internal representations of emotion concepts, which activate in a broad array of contexts which in humans might evoke, or otherwise be associated with, an emotion.
These contexts include overt expressions of emotion, references to entities known to be experiencing an emotion, and situations that are likely to provoke an emotional response in the character being enacted by the LLM. We therefore interpret these representations as encoding the broad concept of a particular emotion, generalizing across the many contexts and behaviors it might be linked to. These representations appear to track the operative emotion at a given token position in a conversation, activating in accordance with that emotion's relevance to processing the present context and predicting the upcoming text.
Interestingly, they do not by themselves persistently track the emotional state of any particular entity, including the AI Assistant character played by the LLM. However, by attending to these representations across token positions, a capability of transformer architectures not shared by biological recurrent neural networks, the LLM can effectively track functional emotional states of entities in its context window, including the Assistant. Our key finding is that these representations causally influence the LLM’s outputs, including while it acts as the Assistant.
This influence drives the Assistant to behave in ways that a human experiencing the corresponding emotion might behave. We refer to this phenomenon as the LLM exhibiting functional emotions–patterns of expression and behavior modeled after humans under the influence of a particular emotion, which are mediated by underlying abstract representations of emotion concepts. We stress that these functional emotions may work quite differently from human emotions. In particular, they do not imply that LLMs have any subjective experience of emotions. Moreover, the mechanisms involved may be quite different from emotional circuitry in the human brain–for instance, we do not find evidence of the Assistant having an emotional state that is instantiated in persistent neural activity (though as noted above, such a state could be tracked in other ways).
Regardless, for the purpose of understanding the model’s behavior, functional emotions and the emotion concepts underlying them appear to be important. The paper is divided into three overarching sections. Part 1 deals with identifying and validating internal emotion-related representations in the model: We extract internal linear representations of emotion concepts (“emotion vectors”) from model activations, using synthetic datasets in which characters experience specified emotions. We validate that these representations activate in scenarios that might be expected to evoke that emotion, and exert causal influence on behavior.
For instance, we demonstrate that when the Assistant is asked to choose between two activities, emotion vector activations evoked by the two choices correlate with, and causally drive, the model’s preference. Part 2 characterizes these emotion vectors in more depth, and identifies other kinds of emotion-related representations at play in the model: The geometry of the emotion vector space roughly mirrors human psychology. Emotions cluster intuitively (fear with anxiety, joy with excitement), and top principal components encode valence (positive vs. negative) and arousal (intensity).
Early-middle layers encode emotional connotations of present content, while middle-late layers encode emotions relevant to predicting upcoming tokens.The representations we find reflect the “operative” emotion in context, rather than tracking a persistent emotional state of a character or speaker. That is, they are locally scoped, encoding the emotional content relevant to processing the context and predicting upcoming text. For example, when a character talks about something dangerous even while otherwise expressing happiness, representations of fear activate.
Note that the “locality” of the representations we find does not preclude the model from tracking characters’ emotional states over long timescales; it can (and does) recall previously cached emotion representations via attention, when they are needed. The model maintains distinct representations for the operative emotion on the present speaker’s versus the other speaker’s turn; these representation…