Anthropic 发表 Toy Models of Superposition 论文,演示神经网络的特征叠加现象
Toy Models of Superposition
Anthropic 发布论文 Toy Models of Superposition,用稀疏合成数据训练的小型 ReLU 网络首次明确演示了神经网络可以在维度不足的情况下以叠加方式表示更多特征。
论文用小型 ReLU 网络直接演示了特征叠加现象,并给出相变和几何结构的理论解释,为理解神经元多义性提供了可参考的框架。
译文尚不完整,完整内容请切换到原文。
如果人工神经网络的单个神经元能够与输入中清晰可解释的特征一一对应,那将非常方便。例如,在一个“理想”的 ImageNet 分类器中,每个神经元只会在特定视觉特征出现时激活,比如红色、左向曲线或狗鼻子。根据经验,在我们研究过的模型中,有些神经元确实能清晰地映射到特征。但特征并不总是如此清晰地对应神经元,尤其是在大型语言模型中,神经元对应清晰特征的情况实际上似乎很罕见。这引出了许多问题。为什么神经元有时与特征对齐,有时却不对齐?为什么有些模型和任务有许多这样的清晰神经元,而在其他模型和任务中它们却极其罕见?
在本文中,我们使用玩具模型——在具有稀疏输入特征的合成数据上训练的小型 ReLU 网络——来研究模型如何以及何时表示比其维度更多的特征。我们将这种现象称为叠加。当特征稀疏时,叠加允许超越线性模型所能做到的压缩,代价是需要非线性过滤的“干扰”。
考虑一个玩具模型,我们在二维空间中训练一个包含五个重要性各异的特征的嵌入其中“重要性”是均方误差损失上的标量乘数。,之后添加一个 ReLU 用于过滤,并改变特征的稀疏性。对于稠密特征,模型学会表示最重要的两个特征的正交基(类似于主成分分析可能给我们的结果),而其他三个特征则不被表示。但如果我们让特征变得稀疏,情况就会改变:
模型不仅可以通过容忍一些干扰以叠加方式存储额外特征,而且我们将表明,至少在某些有限情况下,模型可以在叠加状态下进行计算。(具体来说,我们将表明模型可以将计算绝对值函数的简单电路置于叠加状态。)这使我们假设我们在实践中观察到的 神经网络在某种意义上是在有噪声地模拟更大、高度稀疏的网络。 换句话说,我们训练的模型可能可以被视为在做“与”一个想象中的大得多的模型“相同的事情”,表示完全相同的特征但没有干扰。
特征叠加并不是一个新想法。许多先前的可解释性论文都考虑过它,并且它与数学中研究已久的压缩感知主题密切相关,也与神经科学和深度学习中的分布式、稠密和群体编码思想密切相关。那么,本文的贡献是什么?
对于可解释性研究者而言,我们的主要贡献是直接证明了在相对自然的设置下,人工神经网络中会出现叠加现象,这表明它在实践中也可能发生。也就是说,我们展示了一个案例,其中将神经网络解释为在叠加中具有稀疏结构不仅仅是一种有用的后验解释,而实际上是模型的“真实情况”。我们提供了一个关于这种现象何时以及为何发生的理论,揭示了一个叠加的 相图。这解释了为什么神经元有时是“单义的”,对单个特征做出响应,而有时是“多义的”,对许多不相关的特征做出响应。我们还发现,至少在我们的玩具模型中,叠加表现出复杂的几何结构。
但我们的结果也可能具有更广泛的关注价值。我们发现了初步证据,表明叠加可能与对抗样本和顿悟有关,并且还可能为专家混合模型的性能提供一种理论。更广泛地说,我们研究的玩具模型具有出乎意料的丰富结构,表现出相变、基于均匀多胞形的几何结构、训练过程中类似“能级”的跳跃,以及一种在定性上类似于物理学中分数量子霍尔效应的现象,还有其他一些引人注目的现象。我们最初研究这个主题是为了理解更大模型中可清晰解释的神经元,但我们发现这些玩具模型本身就出乎意料地有趣。
我们玩具模型的关键结果
在我们的玩具模型中,我们能够证明:
- 叠加是一种真实观察到的现象。
- 单义神经元和多义神经元都可以形成。
- 至少某些类型的计算可以在叠加中执行。
- 特征是否以叠加方式存储,由相变所支配。
- 叠加将特征组织成几何结构 ,例如二角形、三角形、五边形和四面体。
我们的玩具模型是简单的 ReLU 网络,因此可以说神经网络至少在某些情况下表现出这些性质,但很难推广到真实网络。
定义与动机:特征、方向和叠加
在我们的工作中,我们经常将神经网络视为具有输入特征,这些特征表示为激活空间中的方向。这不是一个微不足道的论断。我们并不清楚应该期望神经网络表示具有什么样的结构。当我们说“词嵌入具有性别方向”或“视觉模型具有曲线检测神经元”时,我们实际上是在对网络表示的结构做出强有力的断言。
尽管如此,我们相信这种“线性表示假设”既得到了重要的实证发现的支持,也得到了理论论证的支持。可以将其视为两个独立的属性,我们稍后将更详细地探讨:
- 可分解性: 网络表示可以用可独立理解的特征来描述。
- 线性: 特征由方向表示。
如果我们希望逆向工程神经网络,我们需要一种类似可分解性的属性。可分解性正是让我们能够对模型进行推理而无需将整个模型装进脑子里的东西!但仅仅可分解还不够:我们还需要能够以某种方式访问这种分解。为此,我们需要识别表示中的各个特征。在线性表示中,这对应于确定激活空间中的哪些方向对应于输入的哪些独立特征。
有时,识别特征方向非常容易,因为特征似乎对应于神经元。例如,InceptionV1 早期层中的许多神经元明显对应于特征(例如曲线检测器神经元)。为什么我们有时会得到这种极其有用的属性,而在其他情况下却不会?我们假设实际上有两种相互对抗的力量在驱动这一点:
- 特权基:只有某些表示具有特权基,这会促使特征与基方向对齐(即对应于神经元)。
- 叠加:线性表示可以表示比维度更多的特征,使用一种我们称之为叠加的策略。这可以看作神经网络模拟更大的网络。这推动特征远离对应于神经元。
叠加在先前的工作中已被假设,并且在某些情况下,假设类似叠加的东西已被证明有助于找到可解释的结构。然而,我们并不知道此前有特征叠加在神经网络中被明确证明发生(展示了模型叠加这一密切相关现象)。本文的目标就是改变这一点,展示叠加并探索它如何与特权基相互作用。如果叠加发生在网络中,它会深刻影响哪些可解释性研究方法是有意义的,因此明确的证明似乎很重要。
本节的目标将是激发这些想法并详细展开它们。
值得注意的是,本节中的许多想法与其他可解释性研究路线(尤其是解缠)、神经科学(分布式表示、群体编码等)、压缩感知以及许多其他工作路线中的想法有密切联系。本节将专注于阐述我们对这个问题的观点。我们将在相关工作中详细讨论这些其他工作路线。
经验现象
当我们谈论“特征”及其如何被表示时,这最终是围绕若干观察到的经验现象进行的理论构建。在描述我们如何概念化这些结果之前,我们将简单描述一些推动我们思考的主要结果:
- 词嵌入 – Mikolov 等人的一个著名结果发现,词嵌入似乎具有对应于语义属性的方向,允许进行嵌入算术向量,例如
V("king") - V("man") + V("woman") = V("queen")(但参见)。 - 潜在空间 – 类似的“向量算术”和可解释方向的结果也已在生成对抗网络中被发现(例如)。
- 可解释神经元 – 有大量研究结果发现了似乎可解释的神经元(在 RNN 中;在 CNN 中 ;在 GAN 中 ),它们会对某些可理解的属性做出响应。这项工作受到了一些质疑。作为回应,若干论文致力于对少数特定神经元给出极其详细的描述,以期决定性地确立真正检测到某些可理解属性的神经元实例(尤其是 Cammarata 等人 ,但也不止于此)。
- 普遍性 – 在不同网络中都能找到许多对相同属性做出响应的类似神经元。
- 多义神经元 – 与此同时,也有许多神经元似乎并不对输入的可解释属性做出响应,尤其是许多多义神经元 ,它们似乎会对不相关的输入混合做出响应。
因此,我们倾向于认为神经网络表征是由特征组成的,而这些特征被表示为方向。我们将在接下来的章节中展开这一想法。
什么是特征?
我们使用“特征”这一术语,是受我们观察到的神经元(或词嵌入方向)所响应的输入可解释属性所启发。这类被观察到的属性丰富多样!在视觉语境中,这些属性从像曲线检测器和高低频检测器这样的低级神经元,到像定向狗头检测器 或汽车检测器这样更复杂的神经元,再到对应名人、情绪、地理区域以及更多内容 的极其抽象的神经元。在语言模型中,研究人员发现了诸如男性-女性或单数-复数方向这样的词嵌入方向、用于消歧出现在多种语言中的词的低级神经元、更为抽象的神经元,以及有助于生成某些词的“动作”输出神经元。我们希望用“特征”这一术语来涵盖所有这些属性。
但即便有这样的动机,要给出一个令人满意的特征定义也相当具有挑战性。与其提供一个我们确信的单一 定义,我们考虑三种可能的工作定义:
- 特征作为任意函数。 一种方法是将特征定义为输入的任何函数(如 中那样)。但这似乎并不完全符合我们的动机。我们观察到的这些特征有某种特殊之处:它们在某种意义上似乎是用于推理数据的基本抽象,相同的特征会在不同模型中可靠地形成。特征似乎也是可识别的:猫和汽车是两个特征,而猫+汽车和猫-汽车在某种重要意义上似乎是特征的混合,而不是特征。
- 特征作为可解释属性。 我们描述的所有特征对人类来说都出奇地易于理解。人们可以尝试用这一点来定义:特征是输入中人类可理解的“概念”的存在。但允许存在我们可能无法理解的特征似乎也很重要。如果 AlphaFold 发现了某种对预测 蛋白质折叠很重要的化学结构,它很可能并不是我们一开始就能理解的东西!
- 足够大模型中的神经元。 最后一种方法是将特征定义为输入的属性,足够大的神经网络会可靠地专门分配一个神经元来表示它。这个定义比看起来更棘手。具体来说,某物是一个特征,如果存在 一个足够大的模型规模,使得它获得一个专用神经元。这产生了一种类似“epsilon-delta”的定义。我们目前的理解——正如我们将在后续章节中看到的——是任意大的模型仍然可能有很大一部分特征处于叠加态。然而,对于任何给定的特征,假设特征重要性曲线不是平坦的,它最终应该会被分配一个专用神经元。这个定义有助于说明某物是 一个特征——曲线检测器是一个特征,因为你可以在大于某个最小规模的一系列模型中找到它们——但对于我们仅假设或在叠加态中观察到的更常见的特征情况则没有帮助。例如,曲线检测器似乎可靠地出现在足够复杂的视觉模型中,因此是一个特征。对于目前仅在多义神经元中观察到的可解释属性,希望足够大的模型会为它们分配一个专用神经元。这个定义略有循环,但避免了早期定义的问题。
我们撰写这篇论文时,心中秉持的是最终的“足够大模型中的神经元”定义。但我们并不过分执着于它,实际上认为不 prematurely 执着于一个定义可能很重要。Lakatos 的一本著名著作 说明了定义不确定性的重要性,以及在研究背景下重新思考定义往往有多么重要。
特征作为方向
正如我们在前面章节中提到的,我们通常认为特征由方向表示。 例如,在词嵌入中,“性别”和“王室”似乎对应于方向,允许像 V("king") - V("man") + V("woman") = V("queen") 这样的算术运算。可解释神经元的例子也是特征作为方向的情况,因为神经元激活的量对应于表示中的一个基方向
如果特征对应于激活空间中的方向,让我们称神经网络表示为线性的。 在线性表示中,每个特征 f_i 都有一个对应的表示方向 W_i。多个特征 f_1, f_2… 以值 x_{f_1}, x_{f_2}… 激活的存在由 x_{f_1}W_{f_1} + x_{f_2}W_{f_2}… 表示。需要明确的是,被表示的特征几乎肯定是输入的非线性函数。只有从特征到激活向量的映射是线性的。注意,某物是否是线性表示取决于你认为什么是特征。
我们不认为神经网络在经验上似乎具有线性表示是巧合。神经网络由线性函数与非线性函数交错构建。在某种意义上,线性函数构成了绝大多数计算(例如,以 FLOPs 衡量)。线性表示是神经网络表示信息的自然格式!具体来说,有三个主要好处:
- 线性表示是某一层可能实现的显然算法的自然输出。 如果设置一个神经元来对特定的权重模板进行模式匹配,那么当刺激与该模板匹配得越好时,它就会越强烈地激活;匹配得越差,激活就越弱。
- 线性表示使特征“可线性访问”。 典型的神经网络层是一个线性函数后接一个非线性函数。如果前一层中的某个特征以线性方式表示,那么下一层中的神经元就可以“选择它”,并使其稳定地兴奋或抑制该神经元。如果某个特征以非线性方式表示,模型就无法在单一步骤中做到这一点。
- 统计效率。 将特征表示为不同的方向,可能使具有线性变换(例如神经网络的权重)的模型实现非局部泛化,从而提高其相对于只能局部泛化的模型的统计效率。这一观点在 Bengio 的一些著作中尤其得到提倡(例如 )。在这篇博客文章中可以找到一个更易理解的论证。
如果你使用多层,就有可能构造非线性表示并从中检索信息(尽管即使这些例子也可以被视为具有更奇特特征的线性表示)。我们在附录中提供了一个例子。然而,我们的直觉是,非线性表示对于神经网络来说通常效率低下。
人们可能会认为线性表示只能存储与其维度一样多的特征,但事实并非如此!我们将看到,我们称之为叠加的现象将允许模型在线性表示中存储更多特征——可能是多得多的特征。
关于这种特征观如何与将特征视为多维流形的概念相协调的讨论,请参见附录“那多维特征呢?”。
特权基与非特权基
即使特征被编码为方向,一个自然的问题仍然是:哪些方向?在某些情况下,考虑基方向似乎是有用的,但在另一些情况下则不然。这是为什么呢?
当研究人员研究词嵌入时,分析基方向是没有意义的。没有理由期望某个基维度会与任何其他可能的方向不同。理解这一点的一种方式是,想象对词嵌入应用某个随机线性变换 M,并对后续权重应用 M^{-1}。这将产生一个相同的模型,但其基维度完全不同。这就是我们所说的非特权基。当然,在没有特权基的情况下研究激活也是可能的,你只需要以某种方式识别出有趣的方向来研究,例如通过取“man”和“woman”之间的差向量,在词嵌入中创建一个性别方向。
但许多神经网络层并非如此。通常,架构的某些特性会使基方向变得特殊,例如应用激活函数。这“打破了对称性”,使这些方向变得特殊,并可能促使特征与基维度对齐。我们称之为特权基,并将基方向称为“神经元”。通常,这些神经元对应于可解释的特征。
从这个角度来看,只有当神经元处于特权基中时,询问它是否可解释才有意义。事实上,我们通常将“神经元”一词保留给处于特权基中的基方向。(参见更长的讨论此处。)
请注意,拥有特权基并不能保证特征会与基对齐——我们将会看到它们往往并非如此!但这只是让这个问题能够成立的最低条件。
叠加假说
即使存在特权基,神经元也常常是“多义的”,会对多个不相关的特征产生响应。对此的一种解释是叠加假说。粗略地说,叠加的思想是神经网络“想要表示比其神经元数量更多的特征”,因此它们利用高维空间的一个性质来模拟一个拥有更多神经元的模型。
数学中的若干结果表明,类似这样的情况可能是合理的:
- 几乎正交的向量。 虽然在 n 维空间中只可能有 n 个正交向量,但在高维空间中可以有 \exp(n) 个“几乎正交”(余弦相似度 <\epsilon)的向量。参见约翰逊–林登施特劳斯引理。
- 压缩感知。 一般来说,如果将向量投影到较低维的空间中,就无法重建原始向量。然而,如果知道原始向量是稀疏的,情况就会改变。在这种情况下,通常可以恢复原始向量。
具体而言,在叠加假说中,特征被表示为神经元输出向量空间中的几乎正交方向。由于特征只是几乎正交,一个特征激活看起来就像其他特征略微激活。容忍这种“噪声”或“干扰”是有代价的。但对于具有高度稀疏特征的神经网络来说,这一代价可能被能够表示更多特征的好处所抵消!(关键在于,稀疏性大大降低了代价,因为稀疏特征很少激活,因而不会相互干扰,而非线性激活函数则创造了滤除少量噪声的机会。)
一种理解方式是,一个小型神经网络或许能够以带噪声的方式“模拟”一个稀疏的更大模型:
尽管我们是在神经元的意义上描述叠加的,它也可以发生在具有非特权基的表示中,例如词嵌入。叠加仅仅意味着特征数量多于维度数量。
总结:特征性质的层级
本节中的这些想法可以按照神经网络表示可能具有的四个逐渐更严格的性质来理解。
- 可分解性: 可分解的神经网络激活可以分解为特征,而这些特征的含义不依赖于其他特征的值。(这一性质最终是最重要的——参见分解在克服维数灾难中的作用。)
- 线性: 特征对应于方向。每个特征 f_i 都有一个对应的表示方向 W_i。多个特征 f_1, f_2… 以值 x_{f_1}, x_{f_2}… 激活,表示为 x_{f_1}W_{f_1} + x_{f_2}W_{f_2}....
- 叠加与非叠加: 如果 W^TW 不可逆,则线性表示呈现叠加。如果 W^TW 可逆,则不呈现叠加。
- 基对齐: 如果所有 W_i 都是独热基向量,则一个表示是基对齐的。如果所有 W_i 都是稀疏的,则一个表示是部分基对齐的。 这需要一个特权基。
前两个(可分解性和线性)是我们假设普遍存在的性质,而后者(非叠加和基对齐)是我们认为仅有时出现的性质。
演示叠加
如果认真对待叠加假设,一个自然的第一个问题是神经网络是否真的能够以噪声方式表示比其神经元更多的特征。如果不能,叠加假设就可以被轻松否定。
从线性模型得到的直觉会认为这是不可能的:线性模型能做到的最好就是存储主成分。但我们将看到,仅仅添加一点非线性就能使模型以完全不同的方式表现!这将是我们对叠加的首次演示。(这也将是对即使非常简单的神经网络复杂性的一个实例教训。)
实验设置
我们的目标是探索神经网络是否能够将高维向量 x \in R^n 投影到低维向量 h\in R^m,然后恢复它。这个实验设置也可以被视为一个重建 x 的自编码器。
特征向量 (x)
我们首先描述高维向量 x:我们理想化的、解缠结的更大模型的激活值。我们将每个元素 x_i 称为一个“特征”,因为我们想象特征与假设的更大模型中的神经元完全对齐。在视觉模型中,这可能是一个 Gabor 滤波器、一个曲线检测器或一个垂耳检测器。 在语言模型中,它可能对应于一个指代特定名人的 token,或者一个作为特定描述类型的子句。
由于我们没有任何关于特征的真实标签,我们需要为 x 创建合成数据,从建模特征的角度模拟我们认为特征具有的任何重要性质。我们做出三个主要假设:
- 特征稀疏性: 在自然世界中,许多特征似乎是稀疏的,因为它们很少出现。例如,在视觉中,图像中的大多数位置不包含水平边缘、曲线或狗头。在语言中,大多数 token 不指代 Martin Luther King,或者不是描述音乐的子句的一部分。这个想法可以追溯到关于视觉和自然图像统计的经典工作(参见例如 Olshausen, 1997, “Why Sparseness?” 一节)。因此,我们将为特征选择一个稀疏分布。
- 特征比神经元多: 模型可能表示的有用特征数量极其庞大。一个足够通用的视觉模型可能受益于表示它可能看到的每一种动植物以及每一种人造物体。一个语言模型可能受益于表示每一个曾在文字中被提及过的人。这些仅仅触及了可能特征的表层,但似乎已经比任何模型的神经元都要多了。事实上,大型语言模型确实表现出知道那些知名度非常低的人——大概这样的人比它们的神经元数量还多。这一点在神经科学中关于“祖母神经元”合理性的讨论中是一个常见的论点,但对于人工神经网络来说似乎更为有力。真实模型中特征与神经元之间的这种不平衡,似乎必然是神经网络表示中的一个核心矛盾。
- 特征的重要性各不相同: 并非所有特征对给定任务都同样有用。有些特征能比其他特征更多地降低损失。对于一个 ImageNet 模型来说,区分不同品种的狗是一项核心任务,那么一个“垂耳检测器”可能是它能拥有的最重要的特征之一。相比之下,另一个特征可能只能非常轻微地提升性能。出于计算上的原因,我们在本文中不会重点讨论,但我们常常想象存在无限多的特征,其重要性渐近地趋近于零。
具体来说,我们的合成数据定义如下:输入向量 x 是合成数据,旨在模拟我们认为任务真实底层特征所具有的属性。我们将每个维度 x_i 视为一个“特征”。每个特征都有一个关联的稀疏度 S_i 和重要性 I_i。我们让 x_i=0 的概率为 S_i,否则它在 [0,1] 之间均匀分布。选择让特征均匀分布是任意的。指数分布或幂律分布也会非常自然。在实践中,我们关注所有特征具有相同稀疏度的情况,即 S_i = S。
模型 (x \to x')
我们实际上会考虑两个模型,下面会说明动机。第一个“线性模型”是一个已被充分理解的基线,它不表现出叠加。第二个“ReLU 输出模型”是一个非常简单的模型,它确实表现出叠加。这两个模型仅在最终激活函数上有所不同。
线性模型
h~=~Wx
x'~=~W^Th~+~b
x' ~=~W^TWx ~+~ b
ReLU 输出模型
h~=~Wx
x'~=~\text{ReLU}(W^Th+b)
x' ~=~\text{ReLU}(W^TWx + b)
为什么选择这些模型?
叠加假设表明,高维模型中的每个特征对应于低维空间中的一个方向。这意味着我们可以将降维投影表示为一个线性映射 h=Wx。注意,每一列 W_i 对应于低维空间中表示特征 x_i 的方向。
为了恢复原始向量,我们将使用同一矩阵的转置 W^T。这样做的好处是避免了关于低维空间中哪个方向真正对应于某个特征的任何歧义。它在数学上也似乎相对有原则——回想一下,如果 W 是正交归一的,那么 W^T = W^{-1}。虽然 W 不可能真正正交归一,但我们从压缩感知中得到的直觉是,它将在 Candes & Tao 的意义上“几乎正交归一”,并且经验上确实如此。
我们还添加了一个偏置。这样做的动机之一是,它允许模型将其不表示的特征设置为其期望值。但稍后我们会看到,设置负偏置的能力对于叠加现象很重要,这还有第二个原因——大致来说,它允许模型丢弃少量噪声。
最后一步是是否添加激活函数。事实证明,这对于叠加是否发生至关重要。在真实的神经网络中,当特征实际被模型用于计算时,会有一个激活函数,因此在最后包含一个激活函数似乎是合理的。
损失函数
我们的损失是由上述特征重要性 I_i 加权的均方误差:L = \sum_x \sum_i I_i (x_i - x'_i)^2
基本结果
我们的第一个实验将简单地训练几个具有不同稀疏级别的 ReLU 输出模型,并可视化结果。(我们还将训练一个线性模型——如果优化得足够好,线性模型的解不依赖于稀疏级别。)
主要问题是如何可视化结果。最简单的方法是可视化 W^TW(一个特征×特征矩阵)和 b(一个特征长度向量)。注意,特征按从最重要到最不重要排列,因此结果具有相当好的结构。这里是一个小模型(n=20; ~m=5;)的这种可视化可能看起来的示例,该模型以“预期的线性模型式”方式行为,仅表示与其维度一样多的特征:
但我们真正关心的是这个假设的叠加现象——模型是否通过非正交存储来表示“额外特征”?有没有办法更明确地了解它?嗯,一个问题就是模型学会表示多少特征。对于任何特征,它是否被表示由 ||W_i|| 决定,即其嵌入向量的范数。
我们还想了解给定特征是否与其他特征共享其维度。为此,我们计算 \sum_{j\neq i} (\hat{W_i}\cdot W_j)^2,将所有其他特征投影到 W_i 的方向向量上。如果该特征与其他特征正交,则为 0(下面深蓝色)。另一方面,值 \geq 1 意味着存在一些其他特征的组,它们可以像特征 i 本身一样强烈地激活 W_i!
我们可以用这种方式可视化之前查看的模型:
现在我们有了可视化模型的方法,我们可以开始实际进行实验。我们将从考虑只有几个特征的模型开始(n=20; ~m=5;~ I_i=0.7^i)。这将使视觉上容易看到发生了什么。我们考虑一个线性模型,以及几个在不同特征稀疏级别数据上训练的 ReLU 输出模型:
正如我们的标准直觉所预期的那样,线性模型总是学习最重要的前 m 个特征,类似于学习前几个主成分。ReLU 输出模型在稠密特征(1-S=1.0)上表现相同,但随着稀疏度增加,我们看到了叠加现象的出现。该模型通过让特征之间彼此不正交来表示更多特征。 它从较不重要的特征开始,逐渐影响到最重要的特征。最初,这涉及将它们排列成对跖对,其中一个特征的表示向量恰好是另一个的负值,但我们观察到随着它表示更多特征,它逐渐过渡到其他几何结构。 我们将在后面的章节叠加的几何中进一步讨论特征几何。
对于具有更多特征和隐藏维度的模型,结果在定性上相似。例如,如果我们考虑一个具有 m=20 个隐藏维度和 n=80 个特征的模型 (为了考虑更多特征,重要性增加到 I_i=0.9^i),我们观察到的本质上是上述可视化的缩放版本:
数学理解
在上一节中,我们观察到一个令人惊讶的经验结果:在我们的模型输出中添加 ReLU 允许了一种截然不同的解决方案——叠加 ——这在线性模型中不会出现。
出现这种情况的模型在数学上仍然相当简单。我们能否分析性地理解为什么会出现叠加?就此而言,为什么添加一个非线性会使情况与线性模型如此不同?事实证明,我们可以得到一个相当令人满意的答案,揭示我们的模型受平衡两种竞争力量——特征收益 和 干扰 ——的支配,这将是未来有用的直觉。我们还将发现与化学中著名的汤姆森问题的联系。
让我们从线性情况开始。这在先前的工作中已经被很好地理解了!如果想理解为什么线性模型不表现出叠加,简单的答案是观察到线性模型本质上执行 PCA。但这并不完全令人满意:如果我们暂时抛开所有关于线性函数的知识和直觉,究竟为什么叠加不能发生?
更深入的理解可以来自 Saxe 等人的结果,他们研究了线性神经网络 ——即没有激活函数的神经网络——的学习动力学。这类模型最终是线性函数,但由于它们是多个线性函数的复合,动力学可能相当复杂。他们论文的关键点揭示,神经网络权重可以被视为优化一个简单的闭式解。我们可以调整他们的问题,使其与我们的线性情况更相似,我们让模型为 x' = W^TWx,但保持 x 如 Saxe 中那样高斯分布,揭示以下方程:
Saxe 的结果揭示,在所考虑的模型中,从根本上存在两种相互竞争的力量控制着学习动态。首先,模型可以通过表示更多特征来获得更好的损失(我们将其称为“特征收益”)。但如果它表示的特征数量超过了其能够正交容纳的数量,由于特征之间的“干扰”,它也会得到更差的损失。顺便提一下,将线性模型干扰 \sum_{i\neq j}|W_i \cdot W_J|^2 与压缩感知中的 相干性 概念 \max_{i\neq j}|W_i \cdot W_J| 进行对比是很有趣的。我们可以将它们视为同一向量的 L^2 和 L^\infty 范数。事实上,这使得线性模型表示比其维度更多的特征永远不值得。要证明在线性模型中叠加永远不是最优的,可以求解损失梯度为零的条件,或查阅 Saxe 等人的工作。
我们能否对 ReLU 输出模型实现类似的理解?具体来说,我们希望理解 L=\int_x ||I(x-\text{ReLU}(W^TWx+b))||^2 d\textbf{p}(x) 其中 x 的分布使得 x_i=0 的概率为 S。
对 x 的积分根据 ((1\!-\!S)+S)^n 的二项式展开分解为每个稀疏模式对应的项。我们可以将稀疏度相同的项分组,将损失重写为 L = (1\!-\!S)^n L_n +\ldots+ (1\!-\!S)S^{n-1} L_1+ S^n L_0,其中每个 L_k 对应输入为 k-稀疏向量时的损失。注意当 S\to 1 时,L_1 和 L_0 占主导地位。L_0 项对应零向量上的损失,只是对正偏置的惩罚,\sum_i \text{ReLU}(b_i)^2。因此有趣的项是 L_1,即 1-稀疏向量上的损失:
这个新方程与化学中著名的 Thomson 问题 有些相似。特别地,如果我们假设重要性均匀,且固定数量的特征具有 ||W_i|| = 1,其余特征具有 ||W_i|| = 0,并且 b_i = 0,那么特征收益项是常数,干扰项变成一个广义 Thomson 问题——我们只是在球面上以略微不寻常的能量函数打包点。(我们将在后续章节恢复实证研究时看到,这可能是一个富有成效的类比!)
另一个有趣的性质是,ReLU 在 1-稀疏情况下使负干扰免费。这解释了为什么我们看到的解在可能的情况下倾向于只具有负干扰。此外,使用负偏置可以将小的正干扰本质上转化为负干扰。
那么对应于更不稀疏的向量的项呢?我们将这些项的明确写出留给读者,但主要思想是存在多个复合干扰,并且“活跃特征”可能经历干扰。在后面的章节中,我们将看到特征通常组织成稀疏干扰图,使得只有少数特征与另一个特征发生干扰——有趣的是,这降低了复合干扰的概率,并使 1-稀疏损失项相对于其他项更加重要。
叠加作为一种相变
上一节的结果似乎表明,当我们训练模型时,一个特征有三种结果:(1) 该特征可能根本没有被学习;(2) 该特征可能被学习,并以叠加态表示;或者 (3) 模型可能用一个专用维度来表示该特征。这三种结果之间的转变似乎是突变的。可能存在着某种相变。这里,我们使用“相变”的广义含义,即“不连续变化”,而非更技术性的含义,即在无限系统尺寸极限下出现的不连续性。
更好地理解这一点的一种方法是探索是否存在类似物理学中的“相图”,这可以帮助我们理解一个特征何时预期处于这些状态之一。尽管我们可以在之前的实验中看到一些端倪,但很难真正分离出发生了什么,因为许多特征同时变化,并且可能存在交互效应。因此,我们设计了以下实验来更好地分离这些效应。
作为初始实验,我们考虑具有 2 个特征但只有 1 个隐藏层维度的模型。我们仍然考虑 ReLU 输出模型,\text{ReLU}(W^T W x - b)。第一个特征的重要性为 1.0。在一个轴上,我们将第二个“额外”特征的重要性从 0.1 变化到 10。在另一个轴上,我们将所有特征的稀疏性从 1.0 变化到 0.01。然后我们绘制第二个“额外”特征是否未被学习、以叠加态学习,还是被学习并以正交方式表示。为了减少噪声,我们对每个点训练十个模型并对结果取平均,丢弃损失最高的模型。
我们可以将其与一个理论上的“玩具模型的玩具模型”进行比较,在该模型中,我们可以得到不同权重配置的损失作为重要性和稀疏性的函数的闭式解。在一维中存储 2 个特征有三种自然方式:W=[1,0](忽略 [0,1],丢弃额外特征),W=[0,1](忽略 [1,0],丢弃第一个特征以给额外特征一个专用维度),以及 W=[1,-1](以叠加态存储特征,失去表示 [1,1] 的能力,即同时表示两个特征的组合)。我们称最后一种解决方案为“对跖”,因为两个基向量 [1, 0] 和 [0, 1] 被映射到相反的方向。事实证明,我们可以解析地确定这些解决方案的损失(详情可在此 notebook 中找到)。
正如预期,稀疏性是叠加态发生的必要条件,但我们可以看到它与相对特征重要性以一种有趣的方式相互作用。但最有趣的是,似乎存在一个真正的相变,在经验图和理论图中都观察到了!最优权重配置在幅度和叠加态上不连续地变化。(在理论模型中,我们可以解析地确认存在一阶相变:函数之间存在交叉,导致最优损失的导数不连续。)
我们可以对将三个特征嵌入二维空间提出同样的问题。这个问题仍然有一个单一的“额外特征”(现在是第三个),我们可以研究它,询问当我们改变它相对于其他两个特征的重要性并改变稀疏性时会发生什么。
对于理论模型,我们现在考虑四种自然的解决方案。我们可以通过问“W 忽略了哪个特征方向?”来描述这些解决方案。例如,W 可能只是不表示额外的特征——我们将其写作 W \perp [0, 0, 1]。或者 W 可能忽略其他特征之一,W \perp [1, 0, 0]。但有趣的是,有两种方式可以利用叠加来形成对跖对。我们可以将“额外特征”与其中一个其他特征放入对跖对(W \perp [0, 1, 1]),或者将另外两个特征放入叠加,并给额外特征一个专用维度(W \perp [1, 1, 0])。这些解决方案的闭式损失详情可在此 notebook 中找到。我们不考虑将所有特征放入联合叠加的最后一种解决方案,W \perp [1, 1, 1]。
这些图表表明,在编码特征的不同策略之间确实存在相变。然而,我们将在下一节中看到,这种初步视角并未捕捉到更为复杂的结构。
叠加的几何
我们已经看到,叠加可以让模型表示额外的特征,并且随着稀疏性的增加,额外特征的数量也会增加。在本节中,我们将更详细地研究这种关系,发现一个意想不到的几何故事:特征似乎会将自己组织成五边形和四面体等几何结构!在某些方面,本节描述的结构似乎“优雅得令人难以置信”,我们认为它很可能至少部分是我们所研究的玩具模型所特有的。但它似乎值得研究,因为如果其中任何内容能推广到真实模型,它可能会为我们理解其表示提供很大的帮助。
我们将从研究均匀叠加开始,其中所有特征都相同:独立、同等重要且同等稀疏。事实证明,均匀叠加与均匀多面体的几何有着惊人的联系!稍后,我们将继续研究非均匀叠加,其中特征并不相同。事实证明,这至少在一定程度上可以理解为均匀叠加的变形。
均匀叠加
如上所述,我们从均匀叠加开始研究,其中所有特征都具有相同的重要性和稀疏性。我们稍后会看到,这种情况具有一些意想不到的结构,但研究它还有一个更基本的理由:它比非均匀情况更容易推理,并且在实验中我们需要担心的变量更少。
我们想了解当我们改变特征稀疏度 S 时会发生什么。由于所有特征同等重要,我们将不失一般性地假设:将所有特征的重要性按相同比例缩放只会按比例缩放损失,而不会改变最优解。每个特征的重要性为 I_i = 1。我们将研究一个具有 n=400 个特征和 m=30 个隐藏维度的模型,但事实证明,特征数量和隐藏维度并不太重要。特别地,事实证明,只要输入特征数量 n 远大于隐藏维度数量,即 n \gg m,输入特征数量 n 就不重要。而且事实证明,只要我们关注的是学习到的特征与隐藏特征的比率,隐藏维度的数量也并不真正重要。将隐藏维度的数量加倍只会使模型学习到的特征数量加倍。
衡量模型已学习到多少特征的一个便捷方法是查看 Frobenius 范数 ||W||_F^2。由于如果某个特征被表示,则 ||W_i||^2\simeq 1;如果未被表示,则 ||W_i||^2\simeq 0,因此这大致就是模型已学会表示的特征数量。方便的是,该范数与基无关,因此在稠密区域 S=0 中仍然表现良好,此时特征基不受任何特殊对待,模型转而用任意方向表示特征。
我们将绘制 D^* = m / ||W||_F^2,可以将其视为“每个特征的维度数”:
令人惊讶的是,我们发现该图在 1 和 1/2 处“粘滞”。(这非常模糊地类似于分数量子霍尔效应——例如参见此图。)为什么会这样?经检查,1/2 的“粘滞点”似乎对应于一种精确的几何排列,其中特征以“对跖对”的形式出现,每个特征恰好是另一个的负值,从而允许将两个特征打包到每个隐藏维度中。看来对跖对如此有效,以至于模型在广泛的稀疏度范围内优先使用它们。
事实证明,对跖对只是冰山一角。隐藏在这条曲线之下的是许多极其具体的特征几何配置。
特征维度
在上一节中,我们看到存在一个粘滞区域,在某种意义上模型具有“每个特征半个维度”。这是模型所表示特征的平均统计属性,但它似乎暗示了一些有趣的东西。我们能否理解某个特定特征获得了“几分之一维度”?
我们将第 i 个特征的维度 D_i 定义为:
D_i ~=~ \frac{||W_i||^2}{\sum_j (\hat{W_i} \cdot W_j)^2}
其中 W_i 是与第 i 个特征相关的权重向量列,而 \hat{W_i} 是该向量的单位版本。
直观上,分子表示给定特征被表示的程度,而分母则是通过将每个特征投影到其维度上来衡量“有多少特征共享它所嵌入的维度”。在对跖情况下,参与对跖对的每个特征的维度为 D = 1 / (1+1) = 1/2,而未学习到的特征的维度为 0。根据经验,当特征在某种意义上被“高效打包”时,所有特征的维度之和似乎等于嵌入维度的数量。
现在我们可以将上面的图按特征逐一分解。这揭示了更多这样的“粘滞点”!为了帮助我们更好地理解这一点,我们将创建一个带有一些附加信息注释的散点图:
- 我们从上一节中的折线图开始。
- 我们在其上叠加一个散点图,表示每个稀疏度级别下模型中每个特征的各个特征维度。
- 特征维度聚集在某些分数处,因此我们为这些分数画线。(事实证明,每个分数对应一种特定的权重几何结构——我们稍后会讨论这一点。)
- 我们用“特征几何图”来可视化几个模型的权重几何结构,其中每个特征是一个节点,边权重基于点积特征嵌入向量的绝对值。因此,如果特征不正交,它们就会相连。
让我们看看生成的图,然后我们会试着弄清楚它向我们展示了什么:
这些点聚集在特定分数处是怎么回事??我们很快就会看到,模型喜欢创建特定的权重几何结构,并在不同配置之间跳跃。
在上一节中,我们发展了一种将叠加视为相变的理论。但此图上介于 0(不学习特征)和 1(将维度专用于特征)之间的一切都是叠加。叠加发生在特征具有分数维度时。也就是说——叠加不仅仅是一件事!
我们如何将其与我们最初对相变的理解联系起来?我们通常认为水只有三种相:冰、水和蒸汽。但这是一种简化:实际上冰有许多相,通常对应不同的晶体结构(例如六方冰与立方冰)。以一种模糊相似的方式,神经网络特征似乎在“叠加”这一大类中也具有许多其他相。
为什么是这些几何结构?
在上图中,我们发现存在对应于以下维度的不同直线:¾(四面体)、⅔(三角形)、½(对跖对)、⅖(五边形)、⅜(方形反棱柱)和 0(未学习特征)。我们相信,如果不是因为基特征在密集区域中与其他方向无法区分,也会有一条 1(为特征专设维度)的线。
其中几种配置可能会作为著名 Thomson 问题 的解而引人注目。(特别是,方形反棱柱 远不如立方体出名,主要因其作为 Thomson 问题解而在分子几何中的作用 而受到关注。)正如我们之前所见,在非常真实的意义上,我们的模型可以被理解为在求解 Thomson 问题的一个广义版本。当我们的模型选择表示一个特征时,该特征被嵌入为 m 维球面上的一个点。
关于正在发生什么的第二条线索是,Thomson 解中存在对应于均匀多面体(例如四面体)的线,但在我们预期会看到非均匀解的地方,似乎出现了分裂的线(例如,对于三角双锥,我们看到的不是 ⅗ 线,而是在三角形的 ⅔ 处和对跖点的 ½ 处同时出现点)。在均匀多面体中,所有顶点具有相同的几何结构,因此如果我们将特征嵌入为它们,每个特征都具有相同的维度。但如果我们将特征嵌入为非均匀多面体,不同特征之间会有或多或少相互干扰。
特别地,许多 Thomson 解可以理解为较小均匀多面体的 tegum 积 (一种通过将两个多面体嵌入正交子空间来构造多面体的操作)。(在之前的特征几何图形可视化中,两个子图不相连当且仅当它们处于不同的 tegum 因子中。)因此,我们应该预期它们的维度实际上对应于底层的因子均匀多面体。
这也可能解释了为什么尽管我们实际上在研究更高维版本的问题,却观察到了 3D Thomson 问题的解。正如许多 3D Thomson 解是 2D 和 1D 解的 tegum 积一样,也许更高维的解通常是 1D、2D 和 3D 解的 tegum 积。
tegum 积中因子的正交性具有有趣的含义。对于叠加而言,这意味着跨 tegum 因子之间不可能存在任何“干扰”。这可能是玩具模型所偏好的:让许多特征同时相互干扰对它来说可能非常糟糕。(参见我们之前的数学分析中的相关讨论。)
附注:多面体与低秩矩阵
在这一点上,值得明确指出,多面体 与 对称、正定、低秩矩阵 (即形如 W^TW 的矩阵)之间存在对应关系。这种对应关系是我们在上一节中看到的结果的基础,并且通常有助于思考叠加问题。
在某些方面,这种对应关系是平凡的。如果有一个形如 W^TW 的秩为 m 的 n\!\times\!n 矩阵,那么 W 是一个 n\!\times\!m 矩阵。我们可以将 W 的列解释为 m 维空间中的 n 个点。开始变得有趣的地方在于,它清楚地表明 W^TW 是由几何结构驱动的。特别地,我们可以看到非对角项是如何由这些点的几何结构驱动的。
换句话说,多面体与叠加策略之间存在精确的对应关系。例如,在二维空间中将三个特征叠加的每一种策略都对应一个三角形,而每一个三角形都对应这样一种策略。从这个角度来看,如果我们有三个同等重要且同等稀疏的特征,最优策略是等边三角形,这似乎并不令人惊讶。
这种对应关系也适用于相反的方向。假设我们有一个形如 W^TW 的秩为 (n\!-\!i) 的矩阵。我们可以通过 W 没有 表示的那些维度来刻画它——也就是说,哪些方向与 W 正交?例如,如果我们有一个 (n\!-\!1) 矩阵,我们可能会问 W 没有表示的是哪一个方向?如果我们假设在不能表示某些向量的约束下,W^TW 会尽可能“接近单位矩阵”,这一点就特别有启发性。
事实上,给定这样一组正交向量,我们可以从 n 个基向量出发,将它们投影到与给定向量正交的空间中,从而构造出一个多胞体。例如,如果我们从三维出发,然后进行投影使得 W \perp (1,1,1),我们就会得到一个三角形。更一般地,令 W \perp (1,1,1,...) 会给我们一个正 n-单纯形。这很有趣,因为它在某种意义上是“最小可能的叠加”。假设特征同等重要且稀疏,那么不去表示的最佳方向就是完全稠密的向量 (1,1,1,...)!
非均匀叠加
到目前为止,本节关注的是均匀叠加的几何结构,其中所有特征都具有同等重要性、同等稀疏性且相互独立。该模型本质上是在求解 Thomson 问题的一个变体。由于所有特征都相同,对应于均匀多面体的解会获得特别低的损失。在本小节中,我们将研究非均匀叠加,其中特征在某种程度上并不均匀。它们可能在重要性和稀疏性上有所不同,或者具有相关性结构,从而使它们不再独立。这会扭曲我们之前看到的均匀几何结构。
In practice, it seems like superposition in real neural networks will be non-uniform, so developing an understanding of it seems important. Unfortunately, we're far from a comprehensive theory of the geometry of non-uniform superposition at this point. As a result, the goal of this section will merely be to highlight some of the more striking phenomena we observe:
- Features varying in importance or sparsity causes smooth deformation of polytopes as the imbalance builds, up until a critical breaking point at which they snap to another polytope.
- Correlated features prefer to be orthogonal, often forming in different tegum factors. As a result, correlated features may form an orthogonal local basis. When they can't be orthogonal, they prefer to be side-by-side. In some cases correlated features merge into a single feature: this hints at some kind of interaction between "superposition-like behavior" and "PCA-like behavior".
- Anti-correlated features prefer to be in the same tegum factor when superposition is necessary. They prefer to have negative interference, ideally being antipodal.
We attempt to illustrate these phenomena with some representative experiments below.
Perturbing a Single Feature
The simplest kind of non-uniform superposition is to vary one feature and leave the others uniform. As an experiment, let's consider an experiment where we represent n=5 features in m=2 dimensions. In the uniform case, with importance I=1 and activation density 1-S=0.05, we get a regular pentagon. But if we vary one point – in this case we'll make it more or less sparse – we see the pentagram stretch to account for the new value. If we make it denser, activating more frequently (yellow) the other features repel from it, giving it more space. On the other hand, if we make it sparser, activating less frequently (blue) it takes less space and other points push towards it.
If we make it sufficiently sparse, there's a phase change, and it collapses from a pentagon to a pair of digons with the sparser point at zero. The phase change corresponds to loss curves corresponding to the two different geometries crossing over. (This observation allows us to directly confirm that it is genuinely a first order phase change.)
To visualize the solutions, we canonicalize them, rotating them to align with each other in a consistent manner.
These results seem to suggest that, at least in some cases, non-uniform superposition can be understood as a deformation of uniform superposition and jumping between uniform superposition configurations rather than a totally different regime. Since uniform superposition has a lot of understandable structure, but real world superposition is almost certainly non-uniform, this seems very promising!
The reason pentagonal solutions are not on the unit circle is because models reduce the effect of positive interference, setting a slight negative bias to cut off noise and setting their weights to ||W_i|| = 1 / (1-b_i) to compensate. Distance from the unit circle can be interpreted as primarily driven by the amount of positive interference.
A note for reimplementations: optimizing with a two-dimensional hidden space makes this easier to study, but the actual optimization process turns out to be really challenging from gradient descent – a lot harder than even just having three dimensions. Getting clean results required fitting each model multiple times and taking the solution with the lowest loss. However, there's a silver lining to this: visualizing the sub-optimal solutions on a scatter plot as above allows us to see the loss curves for different geometries and gain greater insight into the phase change.
Correlated and Anticorrelated Features
A more complicated form of non-uniform superposition occurs when there are correlations between features. This seems essential for understanding superposition in the real world, where many features are correlated or anti-correlated.
For example, one very pragmatic question to ask is whether we should expect polysemantic neurons to group the same features together across models. If the groupings were random, you could use this to detect polysemantic neurons, by comparing across models! However, we'll see that correlational structure strongly influences which features are grouped together in superposition.
The behavior seems to be quite nuanced, with a kind of "order of preferences" for how correlated features behave in superposition. The model ideally represents correlated features orthogonally, in separate tegum factors with no interactions between them. When that fails, it prefers to arrange them so that they're as close together as possible – it prefers positive interference between correlated features over negative interference. Finally, when there isn't enough space to represent all the correlated features, it will collapse them and represent their principal component instead! Conversely, when features are anti-correlated, models prefer to have them interfere, especially with negative interference. We'll demonstrate this with a few experiments below.
Setup for Exploring Correlated and Anticorrelated Features
Throughout this section we'll refer to "correlated feature sets" and "anticorrelated feature sets".
Correlated Feature Sets. Our correlated feature sets can be thought of as "bundles" of co-occurring features. One can imagine a highly idealized version of what might happen in an image classifier: there could be a bundle of features used to identify animals (fur, ears, eyes) and another bundle used to identify buildings (corners, windows, doors). Features from one of these bundles are likely to appear together. Mathematically, we represent this by linking the choice of whether all the features in a correlated feature set are zero or not together. Recall that we originally defined our synthetic distribution to have features be zero with probability S and otherwise uniformly distributed between [0,1]. We simply have the same sample determine whether they're zero.
Anticorrelated Feature Sets. One could also imagine anticorrelated features which are extremely unlikely to occur together. To simulate these, we'll have anticorrelated feature sets where only one feature in the set can be active at a time. To simulate this, we'll have the feature set be entirely zero with probability S, but then only have one randomly selected feature in the set be uniformly sampled from [0,1] if it's active, with the others being zero.
Organization of Correlated and Anticorrelated Features
For our initial investigation, we simply train a number of small toy models with correlated and anti-correlated features and observe what happens. To make this easy to study, we limit ourselves to the m=2 case where we can explicitly visualize the weights as points in 2D space. In general, such solutions can be understood as a collection of points on a unit circle. To make solutions easy to compare, we rotate and flip solutions to align with each other.
Local Almost-Orthogonal Bases
It turns out that the tendency of models to arrange correlated features to be orthogonal is actually quite a strong phenomenon. In particular, for larger models, it seems to generate a kind of "local almost-orthogonal basis" where, even though the model as a whole is in superposition, the correlated feature sets considered in isolation are (nearly) orthogonal and can be understood as having very little superposition.
To investigate this, we train a larger model with two sets of correlated features and visualize W^TW.
If this result holds in real neural networks, it suggests we might be able to make a kind of "local non-superposition" assumption, where for certain sub-distributions we can assume that the activating features are not in superposition. This could be a powerful result, allowing us to confidently use methods such as PCA which might not be principled to generally use in the context of superposition.
Collapsing of Correlated Features
One of the most interesting properties is that there seems to be a trade off with Principal Components Analysis (PCA) and superposition. If there are two correlated features a and b, but the model only has capacity to represent one, the model will represent their principal component (a+b)/\sqrt{2}, a sparse variable that has more impact on the loss than either individually, and ignore the second principal component (a-b)/\sqrt{2}.
As an experiment, we consider six features, organized into three sets of correlated pairs. Features in each correlated pair are represented by a given color (red, green, and blue). The correlation is created by having both features always activate together – they're either both zero or neither zero. (The exact non-zero values they take when they activate is uncorrelated.)
As we vary the sparsity of the features, we find that in the very sparse regime, we observe superposition as expected, with features arranged in a hexagon and correlated features side-by-side. As we decrease sparsity, the features progressively "collapse" into their principal components. In very dense regimes, the solution becomes equivalent to PCA.
These results seem to hint that PCA and superposition are in some sense complementary strategies which trade off with one another. As features become more correlated, PCA becomes a better strategy. As features become sparser, superposition becomes a better strategy. When features are both sparse and correlated, mixtures of each strategy seem to occur. It would be nice to more deeply understand this space of tradeoffs.
It's also interesting to think about this in the context of continuous equivariant features, such as features which occur in different rotations.
Superposition and Learning Dynamics
The focus of this paper is how superposition contributes to the functioning of fully trained neural networks, but as a brief detour it's interesting to ask how our toy models – and the resulting superposition – evolve over the course of training.
There are several reasons why these models seem like a particularly interesting case for studying learning dynamics. Firstly, unlike most neural networks, the fully trained models converge to a simple but non-trivial structure that rhymes with an emerging thread of evidence that neural network learning dynamics might have geometric weight structure that we can understand. One might hope that understanding the final structure would make it easier for us to understand the evolution over training. Secondly, superposition hints at surprisingly discrete structure (regular polytopes of all things!). We'll find that the underlying learning dynamics are also surprisingly discrete, continuing an emerging trend of evidence that neural network learning might be less continuous than it seems. Finally, since superposition has significant implications for interpretability, it would be nice to understand how it emerges over training – should we expect models to use superposition early on, or is it something that only emerges later in training, as models struggle to fit more features in?
Unfortunately, we aren't able to give these questions the detailed investigation they deserve within the scope of this paper. Instead, we'll limit ourselves to a couple particularly striking phenomena we've noticed, leaving more detailed investigation for future work.
Phenomenon 1: Discrete "Energy Level" Jumps
Perhaps the most striking phenomenon we've noticed is that the learning dynamics of toy models with large numbers of features appear to be dominated by "energy level jumps" where features jump between different feature dimensionalities. (Recall that a feature's dimensionality is the fraction of a dimension dedicated to representing a feature.)
Let's consider the problem setup we studied when investigating the geometry of uniform superposition in the previous section, where we have a large number of features of equal importance and sparsity. As we saw previously, the features ultimately arrange themselves into a small number of polytopes with fractional dimensionalities.
A natural question to ask is what happens to these feature dimensionalities over the course of training. Let's pick one model where all the features converge into digons and observe. In the first plot, each colored line corresponds to the dimensionality of a single feature. The second plot shows how the loss curve changes over the same duration.
Note how the dimensionality of some features "jump" between different values and swap places. As this happens, the loss curve also undergoes a sudden drop (a very small one at the first jump, and a larger one at the second jump).
These results make us suspect that seemingly smooth decreases of the loss curve in larger models are in fact composed of many small jumps of features between different configurations. (For similar results of sudden mechanistic changes, see Olsson et al.'s induction head phase change , and Nanda and Lieberum's results on phase changes in modular arithmetic . More broadly, consider the phenomenon of grokking .)
Phenomenon 2: Learning as Geometric Transformations
Many of our toy model solutions can be understood as corresponding to geometric structures. This is especially easy to see and study when there are only m=3 hidden dimensions, since we can just directly visualize the feature embeddings as points in 3D space forming a polyhedron.
It turns out that, at least in some cases, the learning dynamics leading to these structures can be understood as a sequence of simple, independent geometric transformations!
One particularly interesting example of this phenomenon occurs in the context of correlated features, as studied in the previous section. Consider the problem of representing n=6 features in superposition within m=3 dimensions. If we have the 6 features be 2 sets of 3 correlated features, we observe a really interesting pattern. The learning proceeds in distinct regimes which are visible in the loss curve, with each regime corresponding to a distinct geometric transformation:
(Although the last solution – an octahedron with features from different correlated sets arranged in antipodal pairs – seems to be a strong attractor, the learning trajectory visualized above appears to be one of a few different learning trajectories that attract the model. The different trajectories vary at step C: sometimes the model gets pulled directly into the antiprism configuration from the start or organizes features into antipodal pairs. Presumably this depends on which feature geometry the model is closest to when step B ends.)
The learning dynamics we observe here seem directly related to previous findings on simple models. found that two-layer neural networks, in early stages of training, tend to learn a linear approximation to a problem. Although the technicalities of our data generation process do not precisely match the hypotheses of their theorem, it seems likely that the same basic mechanism is at work. In our case, we see the toy network learns a linear PCA solution before moving to a better nonlinear solution. A second related finding comes from , who looked at hierarchical sets of features, with a data generation process similar to the one we consider. They find empirically that certain networks (nonlinear and deep linear) “split” embedding vectors in a manner very much like what we observed. They also provide a theoretical analysis in terms of the underlying dynamical system. A key difference is that they focus on the topology—the branching structure of the emerging feature representations—rather than the geometry. Despite this difference, it seems likely that their analysis could be generalized to our case.
Relationship to Adversarial Robustness
Although we're most interested in the implications of superposition for interpretability, there appears to be a connection to adversarial examples. If one gives it a little thought, this connection can actually be quite intuitive.
In a model without superposition, the end-to-end weights for the first feature are:
(W^TW)_0 ~~=~~ (1,~ 0,~ 0,~ 0,~ ...)
But in a model with superposition, it's something like:
(W^TW)_0 ~~=~~ (1,~ \epsilon,~ -\epsilon,~ \epsilon,~ ...)
The \epsilon entries (which are solely an artifact of superposition "interference") create an obvious way for an adversary to attack the most important feature. Note that this may remain true even in the infinite data limit: the optimal behavior of the model fit to sparse infinite data is to use superposition to represent more features, leaving it vulnerable to attack.
To test this, we generated L2 adversarial examples (allowing a max L2 attack norm of 0.1 of the average input norm). We originally generated attacks with gradient descent, but found that for extremely sparse examples where ReLU neurons are in the zero regime 99% of the time, attacks were difficult, effectively due to gradient masking . Instead, we found it worked better to analytically derive adversarial attacks by considering the optimal L2 attacks for each feature (\lambda (W^TW)_i / ||(W^TW)_i||_2) and taking the one of these attacks which most harms model performance.
We find that vulnerability to adversarial examples sharply increases as superposition forms (increasing by >3×), and that the level of vulnerability closely tracks the number of features per dimension (the reciprocal of feature dimensionality).
We're hesitant to speculate about the extent to which superposition is responsible for adversarial examples in practice. There are compelling theories for why adversarial examples occur without reference to superposition (e.g. ). But it is interesting to note that if one wanted to try to argue for a "superposition maximalist stance", it does seem like many interesting phenomena related to adversarial examples can be predicted from superposition. As seen above, superposition can be used to explain why adversarial examples exist. It also predicts that adversarially robust models would have worse performance, since making models robust would require giving up superposition and representing less features. It predicts that more adversarially robust models might be more interpretable (see e.g. ). Finally, it could arguably predict that adversarial examples transfer (see e.g. ) if the arrangement of features in superposition is heavily influenced by which features are correlated or anti-correlated (see earlier results on this). It might be interesting for future work to see how far the hypothesis that superposition is a significant contributor to adversarial examples can be driven.
In addition to observing that superposition can cause models to be vulnerable to adversarial examples, we briefly experimented with adversarial training to see if the relationship could be used in the other direction to reduce superposition. To keep training reasonably efficient, we used the analytic optimal attack against a random feature. We found that this did reduce superposition, but attacks had to be made unreasonably large (80% input L2 norm) to fully eliminate it, which didn't seem satisfying. Perhaps stronger adversarial attacks would work better. We didn't explore this further since the increased cost and complexity of adversarial training made us want to prioritize other lines of attack on superposition first.
Superposition in a Privileged Basis
So far, we've explored superposition in a model without a privileged basis. We can rotate the hidden activations arbitrarily and, as long as we rotate all the weights, have the exact same model behavior. That is, for any ReLU output model with weights W, we could take an arbitrary orthogonal matrix O and consider the model W' = OW. Since (OW)^T(OW) = W^TW, the result would be an identical model!
Models without a privileged basis are elegant, and can be an interesting analogue for certain neural network representations which don't have a privileged basis – word embeddings, or the transformer residual stream. But we'd also (and perhaps primarily) like to understand neural network representations where there are neurons which do impose a privileged basis, such as transformer MLP layers or conv net neurons.
Our goal in this section is to explore the simplest toy model which gives us a privileged basis. There are at least two ways we could do this: we could add an activation function or apply L1 regularization to the hidden layer. We'll focus on adding an activation function, since the representation we are most interested in understanding is hidden layers with neurons, such as the transformer MLP layer.
This gives us the following "ReLU hidden layer" model:
h~=~\text{ReLU}(Wx) x'~=~\text{ReLU}(W^Th+b)
We'll train this model on the same data as before.
Adding a ReLU to the hidden layer radically changes the model from an interpretability perspective. The key thing is that while W in our previous model was challenging to interpret (recall that we visualized W^TW rather than W), W in the ReLU hidden layer model can be directly interpreted, since it connects features to basis-aligned neurons.
We'll discuss this in much more detail shortly, but here's a comparison of weights resulting from a linear hidden layer model and a ReLU hidden layer model:
Recall that we think of basis elements in the input as "features," and basis elements in the middle layer as "neurons". Thus W is a map from features to neurons.
What we see in the above plot is that the features are aligning with neurons in a structured way! Many of the neurons are simply dedicated to representing a feature! (This is the critical property that justifies why neuron-focused interpretability approaches – such as much of the work in the original Circuits thread – can be effective in some circumstances.)
Let's explore this in more detail.
Visualizing Superposition in Terms of Neurons
Having a privileged basis opens up new possibilities for visualizing our models. As we saw above, we can simply inspect W. We can also make a per-neuron stacked bar plot where, for every neuron, we visualize its weights as a stack of rectangles on top of each other:
- Each column in the stack plot visualizes one column of W.
- Each rectangle represents one weight entry, with height corresponding to the absolute value.
- The color of each rectangle corresponds to the feature it acts on (i.e. which row of W it's in).
- Negative values go below the x-axis.
- The order of the rectangles is not significant.
This stack plot visualization can be nice as models get bigger. It also makes polysemantic neurons obvious: they simply correspond to having more than one weight.
We'll now visualize a ReLU hidden layer toy model with n=10;~ m=5; I^i = 0.75^i and varying feature sparsity levels. We chose a very small model (only 5 neurons) both for ease of visualization, and to circumvent some issues with this toy model we'll discuss below.
However, we found that these small models were harder to optimize. For each model shown, we trained 1000 models and visualized the one with the lowest loss. Although the typical solutions are often similar to the minimal loss solutions shown, selecting the minimal loss solutions reveals even more structure in how features align with neurons. It also reveals that there are ranges of sparsity values where the optimal solution for all models trained on data with that sparsity has the same weight configurations.
The solutions are visualized below, both visualizing the raw W and a neuron stacked bar plot. We color features in the stacked bar plot based on whether they're in superposition, and color neurons as being monosemantic or polysemantic depending on whether they store more than one feature. Neuron order was chosen by hand (since it's arbitrary).
The most important thing to pay attention to is how there's a shift from monosemantic to polysemantic neurons as sparsity increases. Monosemantic neurons do exist in some regimes! Polysemantic neurons exist in others. And they can both exist in the same model! Moreover, while it's not quite clear how to formalize this, it looks a great deal like there's a neuron-level phase change, mirroring the feature phase changes we saw earlier.
It's also interesting to examine the structure of the polysemantic solutions, which turn out to be surprisingly structured and neuron-aligned. Features typically correspond to sets of neurons (monosemantic neurons might be seen as the special case where features only correspond to singleton sets). There's also structure in how polysemantic neurons are. They transition from monosemantic, to only representing a few features, to gradually representing more. However, it's unclear how much of this is generalizable to real models.
Limitations of The ReLU Hidden Layer Toy Model Simulating Identity
Unfortunately, the toy model described in this section has a significant weakness, which limits the regimes in which it shows interesting results. The issue is that the model doesn't benefit from the ReLU hidden layer – it has no role except limiting how the model can encode information. If given any chance, the model will circumvent it. For example, given a hidden layer bias, the model will set all the biases to be positive, shifting the neurons into a positive regime where they behave linearly. If one removes the bias, but gives the model enough features, it will simulate a bias by averaging over many features. The model will only use the ReLU activation function if absolutely forced, which is a significant mark against studying this toy model.
We'll introduce a model without this issue in the next section, but wanted to study this model as a simpler case study.
Computation in Superposition
So far, we've shown that neural networks can store sparse features in superposition and then recover them. But we actually believe superposition is more powerful than this – we think that neural networks can perform computation entirely in superposition rather than just using it as storage. This model will also give us a more principled way to study a privileged basis where features align with basis dimensions.
To explore this, we consider a new setup where we imagine our input and output layer to be the layers of our hypothetical disentangled model, but have our hidden layer be a smaller layer we're imagining to be the observed model which might use superposition. We'll then try to compute a simple non-linear function and explore whether it can use superposition to do this. Since the model will have (and need to use) the hidden layer non-linearity, we'll also see features align with a privileged basis.
Specifically, we'll have the model compute y=\text{abs}(x). Absolute value is an appealing function to study because there's a very simple way to compute it with ReLU neurons: \text{abs}(x) = \text{ReLU}(x) + \text{ReLU}(-x). This simple structure will make it easy for us to study the geometry of how the hidden layer is leveraged to do computation.
Since this model needs ReLU to compute absolute value, it doesn't have the issues the model in the previous section had with trying to avoid the activation function.
Experiment Setup
The input feature vector, x, is still sparse, with each feature x_i having probability S_i of being 0. However, since we want to have the model compute absolute value, we need to allow it to take on non-positive values for this to be a non-trivial task. As a result, if it is non-zero, its value is now sampled uniformly from [-1,1]. The target output y is y=\text{abs}(x).
Following the previous section, we'll consider the "ReLU hidden layer" toy model variant, but no longer tie the two weights to be identical:
h = \text{ReLU}(W_1x) y' = \text{ReLU}(W_2h+b)
The loss is still the mean squared error weighted by feature importances I_i as before.
Basic Results
With this model, it's a bit less straightforward to study how individual features get embedded; because of the ReLU on the hidden layer, we can't just study W_2^TW_1. And because W_2 and W_1 are now learned independently, we can't just study columns of W_1. We believe that with some manipulation we could recover much of the simplicity of the earlier model by considering "positive features" and "negative features" independently, but we're going to focus on another perspective instead.
As we saw in the previous section, having a hidden layer activation function means that it makes sense to visualize the weights in terms of neurons. We can visualize W directly or as a neuron stack plot as we did before. We can also visualize it as a graph, which can sometimes be helpful for understanding computation.
Let's look at what happens when we train a model with n=3 features to perform absolute value on m=6 hidden layer neurons. Without superposition, the model needs two hidden layer neurons to implement absolute value on one feature.
The resulting model – modulo a subtle issue about rescaling input and output weightsNote that there's a degree of freedom for the model in learning W_1: We can rescale any hidden unit by scaling its row of W_1 by \alpha, and its column of W_2 by \alpha^{-1}, and arrive at the same model. For consistency in the visualization, we rescale each hidden unit before visualizing so that the largest-magnitude weight to that neuron from W_1 has magnitude 1. – performs absolute value exactly as one might expect. For each input feature x_i, it constructs a "positive side" neuron \text{ReLU}(x_i) and a "negative side" neuron \text{ReLU}(-x_i). It then adds these together to compute absolute value:
Superposition vs Sparsity
We've seen that – as expected – our toy model can learn to implement absolute value. But can it use superposition to compute absolute value for more features? To test this, we train models with n=100 features and m=40 neurons and a feature importance curve I_i = 0.8^i, varying feature sparsity.These specific values were chosen to illustrate the phenomenon we're interested in: the absolute value model learns more easily when there are more neurons, but we wanted to keep the numbers small enough that it could be easily visualized.
A couple of notes on visualization: Since we're primarily interested in understanding superposition and polysemantic neurons, we'll show a stacked weight plot of the absolute values of weights. The features are colored by superposition. To make the diagrams easier to read, neurons are faintly colored based on how polysemantic they are (as judged by eye based on the plots). Neuron order is sorted by the importance of the largest feature.
Much like we saw in the ReLU hidden layer models, these results demonstrate that activation functions, under the right circumstances, create a privileged basis and cause features to align with basis dimensions. In the dense regime, we end up with each neuron representing a single feature, and we can read feature values directly off of neuron activations.
However, once the features become sufficiently sparse, this model, too, uses superposition to represent more features than it has neurons. This result is notable because it demonstrates the ability of neural networks to perform computation even on data that is represented in superposition.One question you might ask is whether we can quantify the ability of superposition to enable extra computation by examining the loss. Unfortunately, we can't easily do this. Superposition occurs when we change the task, making it sparser. As a result, the losses of models with different amounts of superposition are not comparable – they're measuring the loss on different tasks! Remember that the model is required to use the hidden layer ReLU in order to compute an absolute value; gradient descent manages to find solutions that usefully approximate the computation even when each neuron encodes a mix of multiple features.
Focusing on the intermediate sparsity regimes, we find several additional qualitative behaviors that we find fascinatingly reminiscent of behavior that has been observed in real, full-scale neural networks:
To begin, we find that in some regimes, many of the model's neurons will encode pure features, but a subset of them will be highly polysemantic. This is similar to the phase change we saw earlier in the ReLU output model. However, in that case, the phase change was with respect to features, with more important features not being put in superposition. In this experiment, the neurons don't have any intrinsic importance, but we see that the neurons representing the most important features (on the left) tend to be monosemantic.
We find this to bear a suggestive resemblance to some previous work in vision models, which found some layers that contained "mostly pure" feature neurons, but with some neurons representing additional features on a different scale.
We also note that many neurons appear to be associated with a single "primary" feature – encoded by a relatively large weight – coupled with one or more "secondary" features encoded with smaller-magnitude weights to that neuron. If we were to observe the activations of such a neuron over a range of input examples, we would find that the largest activations of that neuron were all or nearly-all associated with the presence of the "primary" feature, but that the lower-magnitude activations were much more polysemantic.
Intriguingly, that description closely matches what researchers have found in previous work on language models – many neurons appear interpretable when we examine their strongest activations over a dataset, but can be shown on further investigation to activate for other meanings or patterns, often at a lower magnitude. While only suggestive, the ability of our toy model to reproduce these qualitative features of larger neural networks offers an exciting hint that these models are illuminating general phenomena.
The Asymmetric Superposition Motif
If neural networks can perform computation in superposition, a natural question is to ask how exactly they're doing so. What does that look like mechanically, in terms of the weights? In this subsection, we'll (mostly) work through one such model and see an interesting motif of asymmetric superposition. (We use the term "motif" in the sense of the original circuit thread, inspired by its use in systems biology .)
The model we're trying to understand is shown below on the left, visualized as a neuron weight stack plot, with features corresponding to colors. The model is only doing a limited amount of superposition, and many of the weights can be understood as simply implementing absolute value in the expected way.
However, there are a few neurons doing something else…
These other neurons implement two instances of asymmetric superposition and inhibition. Each instance consists of two neurons:
One neuron does asymmetric superposition. In normal superposition, one might store features with equal weights (eg. W=[1,-1]) and then have equal output weights (W=[1,1]). In asymmetric superposition, one stores the features with different magnitudes (eg. W=[2,-\frac{1}{2}]) and then has reciprocal output weights (eg. W=[\frac{1}{2}, 2]). This causes one feature to heavily interfere with the other, but avoid the other interfering with the first!
To avoid the consequences of that interference, the model has another neuron heavily inhibit the feature in the case where there would have been positive interference. This essentially converts positive interference (which could greatly increase the loss) into negative interference (which has limited consequences due to the output ReLU).
There are a few other weights this doesn't explain. (We believe they're effectively small conditional biases.) But this asymmetric superposition and inhibition pattern appears to be the primary story.
The Strategic Picture of Superposition
Although superposition is scientifically interesting, much of our interest comes from a pragmatic motivation: we believe that superposition is deeply connected to the challenge of using interpretability to make claims about the safety of AI systems. In particular, it is a clear challenge to the most promising path we see to be able to say that neural networks won't perform certain harmful behaviors or to catch "unknown unknowns" safety problems. This is because superposition is deeply linked to the ability to identify and enumerate over all features in a model, and the ability to enumerate over all features would be a powerful primitive for making claims about model behavior.
We begin this section by describing how "solving superposition" in a certain sense is equivalent to many strong interpretability properties which might be useful for safety. Next, we'll describe three high level strategies one might take to "solving superposition." Finally, we'll describe a few other additional strategic considerations.
Safety, Interpretability, & "Solving Superposition"
We'd like a way to have confidence that models will never do certain behaviors such as "deliberately deceive" or "manipulate." Today, it's unclear how one might show this, but we believe a promising tool would be the ability to identify and enumerate over all features. The ability to have a universal quantifier over the fundamental units of neural network computation is a significant step towards saying that certain types of circuits don't exist.Ultimately we want to say that a model doesn't implement some class of behaviors. Enumerating over all features makes it easy to say a feature doesn't exist (e.g. "there is no 'deceptive behavior' feature") but that isn't quite what we want. We expect models that need to represent the world to represent unsavory behaviors. But it may be possible to build more subtle claims such as "all 'deceptive behavior' features do not participate in circuits X, Y and Z." It also seems like a powerful tool for addressing "unknown unknowns", since it's a way that one can fully cover network behavior, in a sense.
How does this relate to superposition? It turns out that the ability to enumerate over features is deeply intertwined with superposition. One way to see this is to imagine a neural network with a privileged basis and without superposition (like the monosemantic neurons found in early InceptionV1, e.g. ): features would simply correspond to neurons, and you could enumerate over features by enumerating over neurons. Superposition also makes it harder to find interpretable directions in a model without a privileged basis. Without superposition, one could try to do something like the Gram–Schmidt process, progressively identifying interpretable directions and then removing them to make future features easier to identify. But with superposition, one can't simply remove a direction even if one knows that it is a feature direction. The connection also goes the other way: if one has the ability to enumerate over features, one can perform compressed sensing using the feature directions to (with high probability) "unfold" a superposition model's activations into those of a larger, non-superposition model.
For this reason, we'll call any method that gives us the ability to enumerate over features – and equivalently, unfold activations – a "solution to superposition". Any solution is on the table, from creating models that just don't have superposition, to identifying what directions correspond to features after the fact. We'll discuss the space of possibilities shortly.
We've motivated "solving superposition" in terms of feature enumeration, but it's worth noting that it's equivalent to (or necessary for) many other interpretability properties one might care about:
- Decomposing Activation Space. The most fundamental challenge of any interpretability agenda is to defeat the curse of dimensionality. For mechanistic interpretability, this ultimately reduces to whether we can decompose activation space into independently understandable components, analogous to how computer program memory can be decomposed into variables. Identifying features is what allows us to decompose the model in terms of them.
- Describing Activations in Terms of Pure Features. One of the most obvious casualties of superposition is that we can't describe activations in terms of pure features. When features are relatively basis aligned, we can take an activation – say the activations for a dog head in a vision model – and decompose them into individual underlying features, like a floppy ear, short golden fur, and a snout. (See the "semantic dictionary" interface in Building Blocks .) Solving superposition would allow us to do this for every model.
- Understanding Weights (ie. Circuit Analysis). Neural network weights can typically only be understood when they're connecting together understandable features. All the circuit analysis seen in the original circuit thread (see especially ), was fundamentally only possible because the weights connected non-polysemantic neurons. We need to solve superposition for this to work in general.
- Even very basic approaches become perilous with superposition. It isn't just sophisticated approaches to interpretability which are harmed by superposition. Even very basic methods one might consider become unreliable. For example, if one is concerned about language models exhibiting manipulative behavior, one might ask if an input has a significant cosine similarity to the representations of other examples of deceptive behavior. Unfortunately, superposition means that cosine similarity has the potential to be misleading, since unrelated features start to be embedded with positive dot products to each other. However, if we solve superposition, this won't be an issue – either we'll have a model where features align with neurons, or a way to use compressed sensing to lift features to a space where they no longer have positive dot products.
Three Ways Out
At a very high level, there seem to be three potential approaches to resolving superposition:
- Create models without superposition.
- Find an overcomplete basis that describes how features are represented in models with superposition.
- Hybrid approaches in which one changes models, not resolving superposition, but making it easier for a second stage of analysis to find an overcomplete basis that describes it.
Our sense is that all of these approaches are possible if one doesn't care about having a competitive model. For example, we believe it's possible to accomplish any of these for the toy models described in this paper. However, as one starts to consider serious neural networks, let alone modern large language models, all of these approaches begin to look very difficult. We'll outline the challenges we see for each approach in the following sections.
With that said, it's worth highlighting one bright spot before we focus on the challenges. You might have believed that superposition was something you could never fully get rid of, but that doesn't seem to be the case. All our results seem to suggest that superposition and polysemanticity are phases with sharp transitions. That is, there may exist a regime for every model where it has no superposition or polysemanticity. The question is largely whether the cost of getting rid of or otherwise resolving superposition is too high.
Approach 1: Creating Models Without Superposition
It's actually quite easy to get rid of superposition in the toy models described in this paper, albeit at the cost of a higher loss. Simply apply an L1 regularization term to the hidden layer activations (i.e. add \lambda ||h||_1 to the loss). This actually has a nice interpretation in terms of killing features below a certain importance threshold, especially if they're not basis aligned. Generalizing this to real neural networks isn't trivial, but we expect it can be done. (This approach would be similar to work attempting to use sparsity to encourage basis-aligned word embeddings .)
However, it seems likely that models are significantly benefitting from superposition. Roughly, the sparser features are, the more features can be squeezed in per neuron. And many features in language models seem very sparse! For example, language models know about individuals with only modest public presences, such as several of the authors of this paper. Presumably we only occur with frequency significantly less than one in a million tokens. As a result, it may be the case that superposition effectively makes models much bigger.
All of this paints a picture where getting rid of superposition may be fairly achievable, but doing so will have a large performance cost. For a model with a fixed number of neurons, superposition helps – potentially a lot.
But this is only true if the constraint is thought of in terms of neurons. That is, a superposition model with n neurons likely has the same performance as a significantly larger monosemantic model with kn neurons. But neurons aren't the fundamental constraint: flops are. In the most common model architectures, flops and neurons have a strict correspondence, but this doesn't have to be the case and it's much less clear that superposition is optimal in the broader space of possibilities.
One family of models which change the flop-neuron relationship are Mixture of Experts (MoE) models (see review ). The intuition is that most neurons are for specialized circumstances and don't need to activate most of the time. For example, German-specific neurons don't need to activate on French text. Harry Potter neurons don't need to activate on scientific papers. So MoE models organize neurons into blocks or experts, which only activate a small fraction of the time. This effectively allows the model to have k times more neurons for a similar flop budget, given the constraint that only 1/k of the neurons activate in a given example and that they must activate in a block. Put another way, MoE models can recover neuron sparsity as free flops, as long as the sparsity is organized in certain ways.
It's unclear how far this can be pushed, especially given difficult engineering constraints. But there's an obvious lower bound, which is likely too optimistic but is interesting to think about: what if models only expended flops on neuron activations, and recovered the compute of all non-activating neurons? In this world, it seems unlikely that superposition would be optimal: you could always split a polysemantic neuron into dedicated neurons for each feature with the same cost, except for the cases where there would have been interference that hurt the model anyways. Our preliminary investigations comparing various types of superposition in terms of "loss reduction per activation frequency" seem to suggest that superposition is not optimal on these terms, although it asymptotically becomes as good as dedicated feature dimensions. Another way to think of this is that superposition exploits a gap between the sparsity of neurons and the sparsity of the underlying features; MoE eats that same gap, and so we should expect MoE models to have less superposition.
To be clear, MoE models are already well studied, and we don't think this changes the capabilities case for them. (If anything, superposition offers a theory for why MoE models have not proven more effective for capabilities when the case for them seems so initially compelling!) But if one's goal is to create competitive models that don't have superposition, MoE models become interesting to think about. We don't necessarily think that they specifically are the right path forward – our goal here has been to use them as an example of why we think it remains plausible there may be ways to build competitive superposition-free models.
Approach 2: Finding an Overcomplete Basis
The opposite strategy of creating a superposition-free model is to take a regular model, which has superposition, and find an overcomplete basis describing how features are embedded after the fact. This appears to be a relatively standard sparse coding (or dictionary learning) problem, where we want to take the activations of neural network layers and find out which directions correspond to features.More formally, given a matrix H \sim [d,m] ~=~[h_0, h_1, …] of hidden layer activations h \sim [m] sampled over d stimuli, if we believe there are n underlying features, we can try to find matrices A\sim [d,n] and B \sim [n,m] such that A is sparse. This approach has been explored by some prior work .
The advantage of this is that we don't need to worry about whether we're damaging model performance. On the other hand, many other things are harder:
- It’s no longer easy to know how many features you have to enumerate. A monosemantic model represents a feature per neuron, but when finding an overcomplete basis there’s an additional challenge of identifying how many features to use for it.
- Solutions are no longer integrated into the surface computational structure. Neural networks can be understood in terms of their surface structure – neurons, attention heads, etc – and virtual structure that implicitly emerge (e.g. virtual attention heads ). A model described by an overcomplete basis has "virtual neurons": there's a further gap between the surface and virtual structure.
- It's a different, major engineering challenge. Seriously attempting to solve superposition by applying sparse coding to real neural nets suggests a massive sparse coding problem. For truly large language models, one would be starting with something like a millions (neurons) by billions (tokens) matrix and then trying to do an extremely overcomplete factorization, perhaps trying to factor it to be a thousand or more times larger. This is a major engineering challenge which is different from the standard distributed training challenges ML labs are set up for.
- Interference is no longer pushing in your favor. If you try to train models without superposition, interference between features is pushing the training process to have less superposition. If you instead try to decode superposition after the fact, whatever amount of superposition is "baked in" by the training process and you don't have part of the objective pushing in your favor.
Approach 3: Hybrid Approaches
In addition to approaches which address superposition purely at training time, or purely after the fact, it may be possible to take "hybrid approaches" which do a mixture. For example, even if one can't change models without superposition, it may be possible to produce models with less superposition, which are then easier to decode.In particular, it seems like we should expect to be able to reduce superposition at least a little bit with essentially no effect on performance, just by doing something like L1 regularization without any architectural changes. Note that models should have a level of superposition where the derivative of loss with respect to the amount of superposition is zero – otherwise, they'd use more or less superposition. As a result, there should be at least some margin within which we can reduce the amount of superposition without affecting model performance. Alternatively, it may be possible for architecture changes to make finding an overcomplete basis easier or more computationally tractable in large models, separately from trying to reduce superposition.
Additional Considerations
Phase Changes as Cause For Hope. Is totally getting rid of superposition a realistic hope? One could easily imagine a world where it can only be asymptotically reduced, and never fully eliminated. While the results in this paper seem to suggest that superposition is hard to get rid of because it's actually very useful, the upshot of it corresponding to a phase change is that there's a regime where it totally doesn't exist. If we can find a way to push models in the non-superposition regime, it seems likely it can be totally eliminated.
Any superposition-free model would be a powerful tool for research. We believe that most of the research risk is in whether one can make performant superposition free models, rather than whether it's possible to make superposition free models at all. Of course, ultimately, we need to make performant models. But a non-performant superposition free model could still be a very useful research tool for studying superposition in normal models. At present, it's challenging to study superposition in models because we have no ground truth for what the features are. (This is also the reason why the toy models described in this paper can be studied – we do know what the features are!) If we had a superposition-free model, we may be able to use it as a ground truth to study superposition in regular models.
Local bases are not enough. Earlier, when we considered the geometry of non-uniform superposition, we observed that models often form local orthogonal bases, where co-occurring features are orthogonal. This suggests a strategy for locally understanding models on sufficiently narrow sub-distributions. However, if our goal is to eventually make useful statements about the safety of models, we need mechanistic accounts that hold for the full distribution (and off distribution). Local bases seem unlikely to give this to us.
Discussion
To What Extent Does Superposition Exist in Real Models?
Why are we interested in toy models? We believe they are useful proxies for studying the superposition we suspect might exist in real neural networks. But how can we know if they're actually a useful toy model? Our best validation is whether their predictions are consistent with empirical observations regarding polysemanticity. To the best of our knowledge they are. In particular:
- Polysemantic neurons exist. Polysemantic neurons form in our third model, just as they are observed in a wide range of neural networks.
- Neurons are sometimes "cleanly interpretable" and sometimes "polysemantic", often in the same layer. Our third model exhibits both polysemantic and non-polysemantic neurons, often at the same time. This is analogous to how real neural networks often have a mixture of polysemantic and non-polysemantic neurons in the same layer.
- InceptionV1 has more polysemantic neurons in later layers. Empirically, the fraction of neurons which are polysemantic in InceptionV1 increases with depth. One natural explanation is that as features become higher-level the stimuli they detect become rarer and thus sparser (for example, in vision, a high-level floppy ear feature is less common than a low-level Gabor filter's edge). A major prediction of our model is that superposition and polysemanticity increase as sparsity increases.
- Early Transformer MLP neurons are extremely polysemantic. Our experience is that neurons in the first MLP layer in Transformer language models are often extremely polysemantic. If the goal of the first MLP layer is to distinguish between different interpretations of the same token (eg. "die" in English vs German vs Dutch vs Afrikaans), such features would be very sparse and our toy model would predict lots of polysemanticity.
This doesn't mean that everything about our toy model reflects real neural networks. Our intuition is that some of the phenomena we observe (superposition, monosemantic vs polysemantic neurons, perhaps the relationship to adversarial examples) are likely to generalize, while other phenomena (especially the geometry and learning dynamics results) are much more uncertain.
Open Questions
This paper has shown that the superposition hypothesis is true in certain toy models. But if anything, we're left with many more questions about it than we had at the start. In this final section, we review some of the questions which strike us as most important: what do we know, and would we like for future work to clarify?
- Is there a statistical test for catching superposition?
- How can we control whether superposition and polysemanticity occur? Put another way, can we change the phase diagram such that features don't fall into the superposition regime? Pragmatically, this seems like the most important question. L1 regularization of activations, adversarial training, and changing the activation function all seem promising.
- Are there any models of superposition which have a closed-form solution? Saxe et al. demonstrate that it's possible to create nice closed-form solutions for linear neural networks. We made some progress towards this for the n=2; m=1 ReLU output model (and Tom McGrath makes further progress in his comment), but it would be nice to solve this more generally.
- How realistic are these toy models? To what extent do they capture the important properties of real models with respect to superposition? How can we tell?
- Can we estimate the feature importance curve or feature sparsity curve of real models? If one takes our toy models seriously, the most important properties for understanding the problem are the feature importance and sparsity curves. Is there a way we can estimate them for real models? (Likely, this would involve training models of varying sizes or amounts of regularization, observing the loss and neuron sparsities, and trying to infer something.)
- Should we expect superposition to go away if we just scale enough? What assumptions about the feature importance curve and sparsity would need to be true for that to be the case? Alternatively, should we expect superposition to remain a constant fraction of represented features, or even to increase as we scale?
- Are we measuring the maximally principled things? For example, what is the most principled definition of superposition / polysemanticity?
- How important are polysemantic neurons? If X% of the model is interpretable neurons and 1-X% are polysemantic, how much should we believe we understand from understanding the X% interpretable neurons? (See also the "feature packing principle" suggested above.)
- How many features should we expect to be stored in superposition? This was briefly discussed in the previous section. It seems like results from compressed sensing should be able to give us useful upper-bounds, but it would be nice to have a clearer understanding – and perhaps tighter bounds!
- Does the apparent phase change we observe in features/neurons have any connection to phase changes in compressed sensing?
- How does superposition relate to non-robust features? An interesting paper by Gabriel Goh (archive.org backup) explores features in a linear model in terms of the principal components of the data. It focuses on a trade off between "usefulness" and "robustness" in the principal component features, but it seems like one could also relate it to the interpretability of features. How much would this perspective change if one believed the superposition hypothesis – could it be that the useful, non-robust features are an artifact of superposition?
- To what extent can neural networks "do useful computation" on features in superposition? Is the absolute value problem representative of computation in superposition generally, or idiosyncratic? What class of computation is amenable to being performed in superposition? Does it require a sparse structure to the computation?
- How does superposition change if features are not independent? Can superposition pack features more efficiently if they are anti-correlated?
- Can models effectively use nonlinear representations? We suspect models will tend not to use them, but further experimentation could provide good evidence. See the appendix on nonlinear compression. For example investigating the representations used by autoencoders with multi-layer encoders and decoders with really small bottlenecks on random uncorrelated data.
Related Work
Interpretable Features
Our work is inspired by research exploring the features that naturally occur in neural networks. Many models form at least some interpretable features. Word embeddings have semantic directions (see ). There is evidence of interpretable neurons in RNNs (e.g. ), convolutional neural networks (see generally e.g. ; individual neuron families ), and in some limited cases, transformer language models (see detailed discussion in our previous paper). However this work has also found many "polysemantic" neurons which are not interpretable as a single concept .
Superposition
The earliest reference to superposition in artificial neural networks that we're aware of is Arora et al.'s work , who suggest that the word embeddings of words with multiple different word senses may be superpositions of the vectors for the distinct meanings. Arora extend this idea to there being many sparse "atoms of discourse" in superposition, an idea which was generalized to other kinds of embedding vectors and explored in more detail by Goh .
In parallel with this, investigations of individual neurons in models with privileged bases were beginning to grapple with "polysemantic" neurons which respond to unrelated inputs . A natural hypothesis was that these polysemantic neurons are disambiguated by the combined activation of other neurons. This line of thinking eventually became the "superposition hypothesis" for circuits .
Separate from all of this, Cheung et al. explore a slightly different idea one might describe as "model level" superposition: can neural network parameters represent multiple completely independent models? Their investigation is motivated by catastrophic forgetting, but seems quite related to the questions investigated in this paper. Model level superposition can be seen as feature level superposition for highly correlated sets of features, similar to the "almost orthogonal bases" experiment we considered above.
Disentanglement
The goal of learning disentangled representations arises from Bengio et al.'s influential position paper on representation learning : "we would like our representations to disentangle the factors of variation… to learn representations that separate the various explanatory sources." Since then, a literature has developed motivated by this goal, tending to focus on creating generative models which separate out major factors of variation in their latent spaces. This research touches on questions related to superposition, but is also quite different in a number of ways.
Concretely, disentanglement research often explores whether one can train a VAE or GAN where basis dimensions correspond to the major features one might use to describe the problem (e.g. rotation, lighting, gender… as relevant). Early work often focused on semi-supervised approaches where the features were known in advance, but fully unsupervised approaches started to develop around 2016 .
Put another way, the goal of disentanglement might be described as imposing a strong privileged basis on representations which are rotationally invariant by default. This helps get at ways in which the questions of polysemanticity and superposition are a bit different from disentanglement. Consider that when we deal with neurons, rather than embeddings, we have a privileged basis by default. It varies by model, but many neurons just cleanly respond to features. This means that polysemanticity arises as a kind of anomalous behavior, and superposition arises as a hypothesis for explaining it. The question then isn't how to impose a privileged basis, but how to remove superposition as a fundamental problem to accessing features.
Of course, if the superposition hypothesis is true, there are still a number of connections to disentanglement. On the one hand, it seems likely superposition occurs in the latent spaces of generative models, even though that isn't an area we've investigated. If so, it may be that superposition is a major reason why disentanglement is difficult. Superposition may allow generative models to be much more effective than they would otherwise be without. Put another way, disentanglement often assumes a small number of important latent variables to explain the data. There are clearly examples of such variables, like the orientation of objects – but what if a large number of sparse, rare, individually unimportant features are collectively very important? Superposition would be the natural way for models to represent this.A more subtle issue is that GANs and VAEs often assume that their latent space is Gaussianly distributed. Sparse latent variables are very non-Gaussian, but the central limit theorem means that the superposition of many such variables will gradually look more Gaussian. So the latent spaces of some generative models may in fact force models to use superposition! On the other hand, one could imagine ideas from the disentanglement literature being useful in creating architectures that resist superposition by creating an even more strongly privileged basis.
Compressed Sensing
The toy problems we consider are quite similar to the problems considered in the field of compressed sensing, which is also known as compressive sensing and sparse recovery. However, there are some important differences:
- Compressed sensing recovers vectors by solving an optimization problem using general techniques, while our toy model must use a neural network layer. Compressed sensing algorithms are, in principle, much more powerful than our toy model.
- Compressed sensing works using the number of non-zero entries as the measure of sparsity, while we use the probability that each dimension is zero as the sparsity. These are not wholly unrelated: concentration of measure implies that our vectors have a bounded number of non-zero entries with high probability.
- Compressed sensing requires that the embedding matrix (usually called the measurement matrix) have a certain “incoherent” structure such as the restricted isometry property or nullspace property . Our toy model learns the embedding matrix, and will often simply ignore many input dimensions to make others easier to recover.
- Features in our toy model have different "importances", which means the model will often prefer to be able to recover “important” features more accurately, at the cost of not being able to recover “less important” features at all.
In general, our toy model is solving a similar problem using less powerful methods than compressed sensing algorithms, especially because the computational model is so much more restricted (to just a single linear transformation and a non-linearity) compared to the arbitrary computation that might be used by a compressed sensing algorithm.
As a result, compressed sensing lower bounds—which give lower bounds on the dimension of the embedding such that recovery is still possible—can be interpreted as giving an upper bound on the amount of superposition in our toy model. In particular, in various compressed sensing settings, one can recover an n-dimensional k-sparse vector from an m dimensional projection if and only if m = \Omega(k \log (n/k)) . While the connection is not entirely straightforward, we apply one such result to the toy model in the appendix.
At first, this bound appears to allow a number of features that is exponential in m to be packed into the m-dimensional embedding space. However, in our setting, the integer k for which all vectors have at most k non-zero entries is determined by the fixed density parameter S as k = O((1 - S)n). As a result, our bound is actually m = \Omega(-n (1 - S) \log(1 - S)). Therefore, the number of features is linear in m but modulated by the sparsity. Note that this has a nice information-theoretic interpretation: \log(1 - S) is the surprisal of a given dimension being non-zero, and is multiplied by the expected number of non-zeros. This is good news if we are hoping to eliminate superposition as a phenomenon! However, these bounds also allow for the amount of superposition to increase dramatically with sparsity – hopefully this is an artifact of the techniques in the proofs and not an inherent barrier to reducing or eliminating superposition.
A striking parallel between our toy model and compressed sensing is the existence of phase changes.Note that in the compressed sensing case, the phase transition is in the limit as the number of dimensions becomes large – for finite-dimensional spaces, the transition is fast but not discontinuous. In compressed sensing, if one considers a two-dimensional space defined by the sparsity and dimensionality of the vectors, there are sharp phase changes where the vector can almost surely be recovered in one regime and almost surely not in the other . It isn't immediately obvious how to connect these phase changes in compressed sensing – which apply to recovery of the entire vector, rather than one particular component – to the phase changes we observe in features and neurons. But the parallel is suspicious.
Another interesting line of work has tried to build useful sparse recovery algorithms using neural networks . While we find it useful for analysis purposes to view the toy model as a sparse recovery algorithm, so that we may apply sparse recovery lower bounds, we do not expect that the toy model is useful for the problem of sparse recovery. However, there may be an exciting opportunity to relate our understanding of the phenomenon of superposition to these and other techniques.
Sparse Coding and Dictionary Learning
Sparse Coding studies the problem of finding a sparse representation of dense data. One can think of it as being like compressed sensing, except the matrix projecting sparse vectors into the lower dimensional space is also unknown. This topic goes by many different names including sparse coding (most common in neuroscience), dictionary learning (in computer science), and sparse frame design (in mathematics). For a general introduction, we refer readers to a textbook by Michael Elad .
Classic sparse coding algorithms take an expectation-maximization approach (this includes Olshausen et al's early work , the MOD algorithm , and the k-SVD algorithm ). More recently, new methods based on gradient descent and autoencoders have begun building on these ideas .
From our perspective, sparse coding is interesting because it's probably the most natural mathematical formulation of trying to "solve superposition" by discovering which directions correspond to features.Interestingly, this is the reverse of how sparse coding is typically thought of in neuroscience. Neuroscience often thinks of biological neurons as sparse coding their inputs, whereas we're interested in applying it the opposite direction, to find features in superposition over neurons. But can we actually use these methods to solve superposition in practice? Previous work has attempted to use sparse coding to find sparse structure . More recently, research by Sharkey et al following up on the original publication of this paper has had preliminary success in extracting features out of superposition in toy models using a sparse autoencoder. In general, we're only in the very preliminary investigations of using sparse coding and dictionary learning in this way, but the situation seems quite optimistic. See the section Approach 2: Finding an Overcomplete Basis for more discussion.
Theories of Neural Coding and Representation
Our work explores representations in artificial “neurons”. Neuroscientists study similar questions in biological neurons. There are a variety of theories for how information could be encoded by a group of neurons. At one extreme is a local code, in which every individual stimulus is represented by a separate neuron. At the other extreme is a maximally-dense distributed code, in which the information-theoretic capacity of the population is fully utilized, and every neuron in the population plays a necessary role in representing every input.
One challenge in comparing our work with the neuroscience literature is that a “distributed representation” seems to mean different things. Consider an overly-simplified example of a population of neurons, each taking a binary value of active or inactive, and a stimulus set of sixteen items: four shapes, with four colors (example borrowed from ). A “local code” would be one with a “red triangle” neuron, a “red square” neuron, and so on. In what sense could the representation be made more “distributed”? One sense is by representing independent features separately — e.g. four “shape” neurons and four “color” neurons. A second sense is by representing more items than neurons — i.e. using a binary code over four neurons to encode 2^4 = 16 stimuli. In our framework, these senses correspond to decomposability (representing stimuli as compositions of independent features) and superposition (representing more features than neurons, at cost of interference if features co-occur).
Decomposability doesn’t necessarily mean each feature gets its own neuron. Instead, it could be that each feature corresponds to a “direction in activation-space”We haven’t encountered a specific term in the distributed coding literature that corresponds to this hypothesis specifically, although the idea of a “direction in activation-space” is common in the literature, which may be due to ignorance on our part. We call this hypothesis linearity., given scalar “activations” (which in biological neurons would be firing rate). Then, only if there is a privileged basis, “feature neurons” are incentivized to develop. In biological neurons, metabolic considerations are often hypothesized to induce a privileged basis, and thus a “sparse code”. This would be expected if the nervous system’s energy expenditure increases linearly or sublinearly with firing rate.Experimental evidence seems to support this Additionally, neurons are the units by which biological neural networks can implement non-linear transformations, so if a feature needs to be non-linearly transformed, a “feature neuron” is a good way to achieve that.
Any decomposable linear code that uses orthogonal feature vectors is functionally equivalent from the viewpoint of a linear readout. So, a code can both be “maximally distributed” — in the sense that every neuron participates in representing every input, making each neuron extremely polysemantic — and also have no more features than it has dimensions. In this conception, it’s clear that a code can be fully “distributed” and also have no superposition.
A notable difference between our work, and the neuroscience literature we have encountered, is that we consider as a central concept the likelihood that features co-occur with some probability.A related, but different, concept in the neuroscience literature is the “binding problem” in which e.g. a red triangle is a co-occurrence of exactly one shape and exactly one color, which is not a representational challenge, but a binding problem arises if a decomposed code needs to represent simultaneously also a blue square — which shape feature goes with which color feature? Our work does not engage with the binding question, merely treating this as a co-occurrence of “blue”, “red”, “triangle”, and “square”. A “maximally-dense distributed code” makes the most sense in the case where items never co-occur; if the network only needs to represent one item at a time, it can tolerate a very extreme degree of superposition. By contrast, a network that could plausibly need to represent all the items at once can do so without interference between the items if it uses a code with no superposition. One example of high feature co-occurrence could be encoding spatial frequency in a receptive field; these visual neurons need to be able to represent white noise, which has energy at all frequencies. An example of limited co-occurrence could be a motor “reach” task to discrete targets, far enough apart that only one can be reached at a time.
One hypothesis in neuroscience is that highly compressed representations might have an important use in long-range communication between brain areas. Under this theory, sparse representations are used within a brain area to do computation, and then are compressed for transmission across a small number of axons. Our experiments with the absolute value toy model shows that networks can do useful computation even under a code with a moderate degree of superposition. This suggests that all neural codes, not just those used for efficient communication, could plausibly be “compressed” to some degree; the regional code might not necessarily need to be decompressed to a fully sparse one.
It's worth noting that the term "distributed representation" is also used in deep learning, and has the same ambiguities of meaning there. Our sense is that some influential early works (e.g. ) may have primarily meant the "independent features are represented independently" decomposability sense, but we believe that other work intends to suggest something similar to what we call superposition.
Additional Connections
After publishing the original version of this paper, a number of readers generously brought to our attention additional connections to prior work. We don't have a sufficiently deep understanding of this work to offer a detailed review, but we offer a brief overview below:
- Vector Symbolic Architectures and Hyperdimensional Computing (see reviews ) are models from theoretical neuroscience of how neural systems can manipulate symbols. Many of the core ideas of how quasi-orthogonal vectors and the "blessings of dimensionality" enable computation are closely related to our notions of superposition.
- Frames (see review ) are a generalization of the idea of a mathematical basis. The way superposition encodes features in lower dimensional spaces might be seen as frames, at least in some cases. In particular, the "Mercedes-Benz Frame" is equivalent to the triangular geometry superposition we sometimes observe.
- Although we discuss compressed sensing and sparse coding above, it's worth noting that this only scratches the surface of research on how sparse vectors can be encoded in lower dimensional dense vectors, and there's a large body of additional work not captured by these topics.
Comments & Replications
Inspired by the original Circuits Thread and Distill's Discussion Article experiment, the authors invited several external researchers who we had previously discussed our preliminary results with to comment on this work. Their comments are included below.
来源:Anthropic:Transformer Circuits · transformer-circuits.pub