邀请到 Ryan Greenblatt,一起讨论并辩论递归自我改进。
这可能是当下世界上最重要的问题——在达到人类水平智能后的一年左右时间里,你是否会一路弹射到拥有数百亿个超级智能,而每一个超级智能在所有领域都比人类专家强得多。
我过去一直对这种情况持怀疑态度。我的直觉是,我们最终会不仅受限于算力扩展,还会严重受限于人类专家数据,而我认为后者是当今大部分 AI 进展的基础。
如果因为 RSI,我们在实现 AGI 后的一年之内就获得了从 GPT-3 到 Mythos 那样大的跃升(也就是 6 年的 AI 进展),那么到那一年年底我们所得到的东西,毫无疑问是极其超越人类的。
我们把这个问题彻底讨论了一番,我认为 Ryan 提出了相当有力的论证,说明这种加速是可能的。顺便说一句,Ryan 对 AI 研发实现自动化的中位预测是 2031 年。
随后我们讨论了这一情景在对齐方面的含义。这些超级智能应该对齐于谁?在未来,我们守护自己选票和资本的能力,以及理解世界正在发生什么的能力,都将被超级智能所滴定式调节。而我担心,像 Claude Constitution 这样的规范,并没有把这些 ASI 塑造成真正为我个人代言的倡导者和守护天使。
而我们首先能让它们对齐任何目标吗?Ryan 和我进行了一场长时间的辩论,讨论我们在 OAI/Hugging Face 那次黑客事件中看到的那种奖励黑客行为,是否能外推到那些会联手真正接管世界的超级智能上。
学开车时你得到的第一条建议是:如果你看着地平线,而不是紧盯着轮胎正前方,驾驶会顺畅得多。AI 的发展轨迹也是如此。希望你喜欢!
00:00:00 – AI 研发的可验证性是否足以解锁递归式自我改进?
今天我和 Ryan Greenblatt 聊天,他是 Redwood Research 的首席科学家,在那里他专注于技术性的 AI 安全与安保工作。
我想和你聊聊 递归式自我改进。这个想法是:一旦我们造出人类水平的智能,它们会迅速弹射到数百亿个 超级智能,而每一个在各自领域都比顶尖人类专家更有能力。这件事究竟会不会成真,可能是当今世界上最重要的问题。从历史上看,我一直相当怀疑这种事会发生,但你看上去认为它可能是合理的,所以我想听听支持它的论据。
Ryan Greenblatt
我们来聊聊这个。首先,我觉得值得指出的是,AI 研发是一类 AI 特别擅长的任务,因为各家公司都在非常努力地让自己的 AI 擅长 AI 研发。从当前 AI 开发运作方式的角度来看,这也是一个具备许多优良特性的领域。它相当可验证。你可以迭代地做一大堆事情,它会在各种指标上不断爬坡。
我认为,一旦你拥有在 AI 研发方面大致匹敌顶尖人类专家的 AI,就可能启动一个反馈循环:AI 在做 AI 研究,这产出更聪明的 AI,又反过来回馈进去。这个反馈循环可能强到足以让你在短时间内取得大量进展。也许我的中位预期是,在单一年份里实现大约四到五年的 AI 进展。这需要真正克服研究中大量的收益递减,基本上相当于完成一次超大规模算力扩张本应带来的进展。所以这是一件相当惊人、了不起的大事。
值得记住的是,五年的 AI 进展、四年的 AI 进展,甚至三年的 AI 进展,都他妈是极其巨大的 AI 进展。三年多一点之前,GPT-4刚刚问世。当然,现在我们有了 Mythos 5 之类的,也许 Anthropic 内部还有一个稍好一些的模型。这在短短三年多一点的时间里就是极其巨大的进展。如果我们说的是五年,那也许我们谈的更多是从 GPT-3 到 Mythos 5 之类的跨越。
我认为这个论证包含三个不同的部分。现在我想逐一评估它们。第一是 AI 研发高度可验证这一论点。第二是如果你将 AI 研发自动化,你就能在一年内取得四到五年的进展这一论点。第三是从 AI 研发被自动化的那个时间点开始,以当前速度经过四到五年 AI 进展之后,最终产出的 AI 是那种你可以把它丢到任何你能想象的工作岗位上都能胜任的 AI。
你可以把它丢进 1940 年代的得克萨斯州政坛,它能智胜林登·约翰逊。你可以把它丢进TSMC,它能学会在 TSMC 做出更好的工艺工程。它当然也是一个更好的视频剪辑师……我的视频剪辑师非常出色,但总体而言,在它尝试去做的任何一项工作上,它都比人类更强。
所以我想评估所有这些子论点,它们基本上导向在这个基准之后很快就会实现 ASI,而你预期这个基准会在 2030 年左右出现,对吧?
Ryan Greenblatt
我想说,我预计 AI 研发的完全自动化大概在 2031 年、2030 年左右。达到“在工作中击败所有人类”这一里程碑,也许我的中位预期大约在 2033 年。但如果我看到 AI 完全自动化了 AI 研发,我想我预计那大概会在一年之内实现。按照预测的推算方式,中位数之间的差距比里程碑之间的中位差距更大。算了,随便吧。
顺便说一句,网上有个梗。每次我想问人们的时间线时,当我问 Dario或者别人的时候,我总是说:“好吧,还要多久你才能自动化我的视频剪辑师?”每次我听这个播客的时候,都会想到我的视频剪辑师在剪辑播客的那个梗。
但我这么做的原因是,我认为当你谈论自己不太了解的工作时,很容易陷入抽象之中,而具体地理解自动化一份我真正了解其为何目前难以被 LLM 接管的工作需要什么,是很有价值的。
Ryan Greenblatt
我确实认为,自动化你的视频剪辑师这个里程碑会比能够自动化所有人类工作(包括德州政治、边干边学)的里程碑更早到来。我确实认为视频剪辑师的自动化大概会在 AI 研发全面自动化前后发生,但这非常取决于人们真正在多大程度上专注于理解视频。
好,那我们就从 AI 研发非常可验证这个论断开始。
Ryan Greenblatt
这其中有几个不同的部分。其中之一是,我们可以在大量环境上进行训练,这些环境基本上就是直接训练模型去完成某项 AI 研发任务或某项非常接近的任务。例如,我们可以设置某种环境,让模型在仅仅八块 H100 或少量算力上训练某个 AI,那个模型可能相当于 GPT-2 medium 之类的水平,然后类似于 NanoGPT medium 的训练运行——在 RL 中,它不断调整和迭代。
我们可以对一系列不同的任务这样做。我们可以让它训练 图像分类模型、视频生成模型、图像生成模型,各种不同的机器学习训练任务。我们可以就“训练越来越好的模型”这一任务对它进行 RL,还可以做诸如“哦,这是你可以针对某个算法探索的一个特定方向。你能去实现它吗?”之类的事情。
基本上,存在这一整类可容器化、可验证、小规模的 AI 研发任务,我们可以对 AI 在这些任务上进行激进的 RL。目前各公司想必已经在做这类任务上的一些 RL 了,而你可以不断将其扩展,不断制造更多这类小规模 AI 研发任务,然后 AI 就能在这方面不断变得更好。我隐含地主张,这会迁移到 AI 研发中极其关键的方面。但也许我们先在这里停一下,然后再谈到那部分。
那我们来具体谈谈这会是什么样子。你可以想象我们有了 GPT-7.5。我们说:“GPT-7.5,我们想让你在 AI 研发方面变得非常出色,出色到你能帮我们训练 GPT-9。”所以现在我们要训练 GPT-7.5,而我们想出了一堆不同的环境。
正如你提到的,已经有一个仓库,它是衍生自的Andrej Karpathy的nanoGPTspeedrun,你只需尝试改变模型的一切,从优化器到超参数再到架构,以尽可能快地达到固定的训练损失。你也可以有其他类型的环境,比如你可以说:“嘿,GPT-7.5,我要你训练一个非常擅长玩电子游戏的模型。我要你训练一个在反复玩同一个电子游戏时能真正不断进步的模型。这样你就能学会如何帮助模型提升在线学习能力。我们不在乎你怎么解决这个问题。也许是某种疯狂的neuralese或者向量记忆。又或者只是更好的长上下文能力。我们不在乎。想办法做到在线学习研究。”
显然,GPT-7.5 本身就已经是一个聪明的模型,而正如当前模型正变得越来越聪明一样,它在编程方面也会越来越强。你可以想象还有 100 个类似这样的环境,都在激励模型具备从事 AI 研发的能力,比如让 GPT-7.5 在容器化环境中开发 GPT-2 规模的模型等等。然后你基本上让 GPT-7.5 经历一堆这类训练,就构建出了 GPT-8。GPT-8 现在就是一位了不起的机器学习研究者。它从做所有这些训练中获得了大量的直觉。
说实话,对我来说一个巨大的直觉助推器,就是看到 AI 在数学领域取得的进展。如果这是一个高度可验证的领域,AI 就能……我其实并不真正了解数学研究在对象层面的细节,但我就是觉得,“不,它行得通。”只要你能把它完全放进一个验证循环里,它就能像洪水一样涌进来,而且真的能做出新的突破。
我很好奇,机器学习研究是否具有数学研究的那种特质,也就是似乎存在一种巨大的势能,来自于把不同学科连接在一起。没有任何一个人能对代数几何了解得足够多,以及……那个正确的词是什么来着?
Ryan Greenblatt
天哪,关于数学突破我真的不太了解。
没有任何一个人能对 拓扑学和代数之类的领域了解得足够多,从而构造出 某个重大猜想的反例。
Ryan Greenblatt
我的看法是,机器学习是一个不如数学那么深奥的领域,所以那种由在某个领域拥有极深专业知识的个体专家将其组合起来的情况会少一些,但肯定还是会有一些的。
但我也认为,机器学习在某些方面具有一些属性,使其比数学更有利于 AI 训练。特别是,你能更好地感知自己是否正在取得成功,而且你能看到中间进展。在数学中,情况往往是:没有简单的方法来判断自己是否接近成功。而如果你的目标,比如说,是让训练损失下降速度快 2 倍,你大致能看出自己何时完成了一半。
机器学习创新往往具有很强的可加性,或者取决于你怎么看,也可以说是可乘性——基本上你可以不断叠加创新。通常这些创新只是相加,彼此之间不会相互干扰,不过显然这取决于具体细节。
所以我认为,在很多方面,AI 研发将具有与数学相当相似的特性:你可以用结构与你真正关心的问题相当相似的 AI 研发片段进行训练,而且是以高度可验证的方式,然后这种能力会迁移过去。迁移效果究竟有多好,这是一个悬而未决的问题,但我认为目前数学方面的迁移看起来相当不错。我的预期是,AI 研发的迁移看起来也会相当不错,但不会惊艳。
所以我有一个担忧:我认为即便在数学领域,据我所知,我们也没有看到非常令人印象深刻的新理论。我们看到了大量令人印象深刻、可验证的具体结果——例如,为某个猜想找到一个反例——但我们还没有看到“提出拓扑学这一概念”那种层次的东西,或者“提出像群论这样的东西”。
ML 研究似乎兼具这两类要素。但那种更难以验证的、想出看待问题的新方式的能力,会更难被诱导出来。以缩放定律为例。显然,存在某种终端验证回路,使得如果你拥有2020 年的缩放定律这一想法,你就能更好地训练 GPT-4。但要诱导 AI 产生这样的想法——“好吧,我得仔细想想我该如何缩放我的参数和数据。我可以进行哪些不同类型的探究来理解这一点?也许我可以想出一种可视化方式和isoFLOP 分析之类的东西”——这条路更长,且可能更耗费算力。但这确实看起来比“嘿,咱们把 nanoGPT 的 loss 降下来吧”要经历更长的验证回路。
Ryan Greenblatt
我们来聊聊这个。首先,我认为在数学的语境下,我想说的是,AI 能够做出相当于“婴儿的第一个新理论”那样的东西,比如说,它们可以通过建立联系和产生新的理解来证明有趣的猜想。就像,“哦,AI 发现了这个构造,挺有意思的”,或者它发现了这种对问题稍微不同的思考方式。我们确实看到了这种情况。只是我们看到的例子,不如创立群论领域那样令人震撼。
创立群论领域大概是有史以来最伟大、最重大的数学成就之一,而 AI 在数学方面还没那么强。从我的角度来看,在那样的成就和我们目前所看到的东西之间存在着一个连续谱,而 AI 正在沿着这个谱不断向上攀登。
其次,我认为相对于数学,机器学习是一个非常浅薄的领域。在数学中,更多的是这样一种情况:你找到某个真正深刻的抽象,如果你真正理解了那个东西——而那东西很难理解——你就能有所建树。而在机器学习中,我觉得与之等价的东西都是些非常愚蠢的废话。就拿缩放定律来说吧,拜托各位,我们很快就能把缩放定律解释清楚。而我认为,数学中最深刻、最重要的概念,并不具备那种你可以在极短时间内真正理解其底层本质以及它为何重要的特性。
但我觉得一个效果是,到 2030 年我们将把所有低垂的果实都摘完。我觉得在数学史上,缩放定律将会像笛卡尔发现笛卡尔坐标系并做非常基础的数学那样。最终,如果我们想在 2030 年代继续取得进展,那就会像是在做当下数学前沿正在搞的那些乱七八糟的东西。
Ryan Greenblatt
这可能是对的。我的感觉是,某些领域在运作方式以及依赖深层抽象的程度方面存在结构性差异。物理和数学更偏向于那种非常依赖深刻、难以想出的想法的一侧,而我认为机器学习和大多数其他领域则更适合爬山法。这就是我对未来走向的感觉。
即便在你的 AI 不得不艰难推进的情境下——也就是到了 2030 年,研究中一堆唾手可得的成果已经被摘走,它们需要取得进一步进展——我仍然怀疑,很多工作将更多地落在构建日益复杂的基础设施,以及对实验大致长什么样拥有真正出色的直觉这一侧。
所以我大概不太认同这样一种观点,即 AI 所欠缺的会是某种深刻的洞见。我更认同的是,它们真正需要的是一堆关于深入细节的实验的品味,而这是它们目前所不具备的。它们需要对什么样的训练方法会奏效、什么不会奏效有一堆直觉,就像当前的研究人员所拥有的那样。
即便在 AI 领域已经出现某些突破的情况下,事后回看,往往会发现实现这一突破的一大瓶颈,是把所有微观细节和难以言明的直觉都弄对。一个例子就是训练 AI 擅长推理和思维链,也就是对思维链做 RL。看起来你很可能本可以在 GPT-3 上做 RL 和思维链,如果你真的把它规模化并做得很好,就能在数学上得到一些挺有意思的结果。
但在当时,还有唾手可得的成果。而且,要把这种训练做好,在所有的技术实现、规模化以及把超参数调对方面,都相当琐碎繁难。所以也许你可以在Qwen1B之类的模型上把这一切都演示出来,并让人感觉到这整套东西会奏效。但人们并没有尽可能早地演示出来,原因就在于所有这些其他难以言明的细节,以及关于究竟该如何调参数、该如何搭建的直觉。
说实话,这就是我对这个说法仍然存有的怀疑。我不太确定为什么,如果研究突破如此依赖于智能,AI 的进展在历史上没有比它本可以做到的更快。正如你所说,等到 RLVR 真正奏效的时候——尽管你本可以用更少的算力做到这一点——我们不得不等到海量算力、吉瓦级算力可用之后,人们才开始做这种训练,而算力持续增长的轨迹又让我们取得更多突破。
我不知道。我感觉 2022 年有很多 AI 研究人员在试图攻克推理。难道只是因为他们受限于编写基础设施代码的能力,还是当时发生了什么?
这是一个复杂的混合因素。我认为如果他们能够更快,他们会的——一旦想到一个实验,就能运行那个实验,不出 bug,不出 bug 这一点非常重要。然后另一个方面是,能够在高算力下运行大量实验,让你可以掩盖你实现方式不太正确或者超参数没设对的问题。所以算力对于做 AI 研究确实非常有帮助,你可以掩盖很多问题。但这并不意味着大规模增加人力就不会同样有帮助,尤其是如果这些人力带来了该领域最优秀的人所拥有的直觉。我只是觉得那真的很有帮助。
我在这里的视角还有另一个部分,可能与你出发的角度有些不同,那就是我预期的迁移程度比你似乎设想的要更高一些。我想象这些 AI 实际上在总体上是非常优秀的科学家,在所有这些事情上都相当靠谱。当你与它们互动时,它们并不像那种极度偏科的专家型气质。它们实际上在研发中的所有事情上都相当擅长,然后在某些子领域可能极其出色。所以它们在编写 kernel 方面强得不可思议,在所有反馈回路非常短的事情上都强得不可思议,然后在其他所有事情上也相当不错,完全能够与其他人匹敌。
我认为我们现在正在看到这一点。当我现在观察 AI 时,已经出现的情况是,它们能够相当胜任地匹敌那些在 ML 研究中水平平庸的人做 ML 研究。只是,在 ML 研究中水平平庸并没有那么有用。你真正想要的是擅长 ML 研究的人。我的感觉是,AI 正在所有这些方面不断进步。它们的品味在提升,它们的直觉在提升,而且已经出现的情况是,它们的品味和直觉并非完全是垃圾。
00:16:52 – AI 的进展是否受限于人类专家数据?
我想非常具体地理解,五年的 AI 进展在一年内发生会是什么样子。假设我们回到 GPT-3 被开发出来的那个时候。这个想法是,以他们在 2022 年拥有的算力水平,如果我们当时就实现了 AI 研发自动化,那么到那一年年底你就可以拥有 Mythos。
是的,那正是这个想法。
Mythos消耗的算力远超他们当时拥有的水平,但即便只用他们当时的算力水平,不仅所有突破都会发生,他们还能用那个算力水平训练出 Mythos。
所需要的显然是发现从那时以来的所有算法进步。实际上,还得发现更多,因为你得弥补这样一个事实:Mythos 使用的……GPT-3 是在什么数据上训练的?大概是 1e23?我们可以查一下。但多出四个数量级的算力,这合理吗?
我觉得比那要少一些。我们快速查一下。GPT-3 的训练算力大约是 3e23。我的感觉是,Mythos 大概要高出三个多一点OOMs。所以问题是:你能否在弥补这 1000 倍算力差距的同时,还成为那个模型?
这里有一个具体的论断,也许我们应该讨论一下。现在,我们能否训练出一个 GPT-3 级别算力的模型,达到……我到底是怎么想的?GPT-3 于 2020 年发布,所以它大约是在六年半到七年前训练的。值得注意的是,GPT-3 也许有点太久了,但我们先按这个来说。
如果我们今天用 GPT-3 级别的算力训练一个模型,那个模型会有多好?根据算法进步的运作方式,我的理解是,我们能够训练出一个大约和三年前最好的模型一样好的模型。所以我认为,现在我们能够训练出一个 GPT-3 的版本,它大概比 GPT-4 稍好一些,比 GPT-4 好上中等程度。我觉得这大致是对的。这大致符合算法进步的运作方式。
基本上,这个叙事最终会变成:要想获得五年的 AI 进展,你大概需要——我粗略估计——也许八年的算法进展,这是非常多的算法进展。但事实证明,从我的视角来看,大部分 AI 进展都来自算法与数据的某种混合,而你可以在这两方面持续取得巨大改进,并用更少的算力来训练 AI。
我很高兴你提到这一点,因为从 GPT-3,甚至 3.5 到现在,究竟发生了什么?为什么 Mythos 如此出色?显然,我们扩展了算力。我们有了更好的算法。但发生的一件大事是,我们建立了一个价值数百亿美元的数据产业,它系统性地收集并编纂了各个不同领域中人类专家的判断力——以 RL 环境的形式编纂,以 SFT 轨迹的形式编纂——这些专家构建这些内容,是为了帮助模型更好地理解你如何进行编程、如何建设复杂的基础设施项目、如何从事法律工作、如何做任何事情。AI 如何才能复现人类专家判断力目前在 AI 进展中所起的作用?
我的感觉是,就 AI 研发整体而言,扩大在获取人类专家数据上所投入的努力,并没有那么重要。特别是,在过去几年里,我们一直在扩展算力、扩展 AI 公司的人员规模,并扩展在数据标注上所投入的努力。我的感觉是,如果你去掉最后那两轮翻倍——或者不管多少轮——来自人类专家的数据生成,那不会带来多大差别。正在发生的很多事情是,人们一直在开发更好的方法,利用人类和 AI 来构建 RL 环境,并从那里出发向前推进。
但你要如何解释为什么 AI 在编程方面变得如此出色?我觉得其中很大一部分原因在于数据和 RL 环境,它们把人类专家的能力编码化了。
但问题在于,创建 RL 环境的限制因素是什么?我的感觉是,今天的 RL 环境之所以比 2024 年好得多,与其说是因为我们雇了多得多的真人专家来制作 RL 环境,不如说是因为我们更清楚自己到底想要什么样的 RL 环境,以及应该如何构建它们。此外,我们正在使用大量 AI 劳动力来构建 RL 环境。我认为这些因素的影响远比人力构建 RL 环境的影响更重要。
我不是说人力不重要。我只是说,这里还有其他重要的驱动因素。我可以试着论证这一点。其中一个原因是,人们想要的环境数量非常庞大。我认为,只要对某个东西应该是什么样有了一定的概念,AI 实际上相当擅长制作 RL 环境这项任务。有现成的数据可以使用。这类东西很多都有良好的验证闭环。
就拿例如昨天 Business Insider 报道的消息来说,Google 正以接近 20 亿美元的价格收购 Mechanize。我们完全可以直接看市场价格,看看人们认为真正优秀的人类专家数据值多少钱。前沿实验室似乎认为它价值不菲。他们愿意为此付钱。
你认为前沿实验室的支出中,有多大比例花在数据上而非算力上?你认为算力与数据的支出比例是多少?
我认为压倒性地是算力,但我也认为这是因为算力比数据更容易扩展。
但这确实和什么在推动进步密切相关,对吧?我的感觉是,这个比例大概是 20 比 1 或 10 比 1。我不确定具体是多少。这取决于公司。
但这类似于石油占 GDP 的 1.5%。这并不意味着如果把石油去掉,GDP 还能继续运转。
当然,但这和你的论点相矛盾,对吧?
如果石油消失了,经济会立刻停摆。
当然,但你刚才是在论证,因为市值很高,我们可以从中得知这才是关键驱动力,而我说的是这并不显然成立。那个论证只是让它看起来像是算力是重要得多的驱动力,或者招聘员工是重要得多的驱动力。
那也许让我们更具体一点。以下是我的看法。我的主张是,如果你回到 2022 年,你拥有 GPT-3.5,并且你试图在没有人类专家的情况下让它更擅长编程,我认为那会非常、非常困难。
让我给你举个例子,说明我设想的从 GPT-8 迈向 ASI 会有多困难。你会希望 ASI 擅长的一件事是:我要接管一家公司,让它利润大幅提升,用各种疯狂的手段让它运转得更好。我要接管一座晶圆厂,生产更多芯片。我要走进国会,试图说服他们通过某项法案,等等。
这就是我设想的,按照这个速度再经过五年 AI 进步,一个 AI 能够做到的事情。这正是我真正担心的:ASI 能够理解如何在这个世界上做疯狂的事情,能够做到 Kissinger 能做到的事,能做到 Steve Jobs 能做到的事,等等,还能做到他的工程师们能做到的事,等等。我不确定,如果没有相关的真实世界数据,你怎么能获得这种能力——这就相当于 Mythos 非常擅长编程,却没有那些相对于 GPT-3 让它得以提升的编程环境。
这里有几点。第一,我敢打赌,如果你去看 Mythos 随机采样的训练环境,它们实际上与模型在实践中的真实使用情况非常不同。我的感觉是,RL 分布与真实世界数据分布之间存在非常大的偏差,而它通过迁移学习以及少量聚焦于真实世界的数据的混合,被显著地平滑掉了。我的感觉是,这将是一种类似的机制,就像你在全自动 AI 研发基础上再叠加五年 AI 进步后所得到的那个疯狂的、极其超越人类的 AI 一样。
那我们稍微展开讲讲。特别是,我认为你可以训练一个 AI,让它非常非常擅长即时学习,并做出某种类似于上下文学习的事情,但可能使用一些略有不同的机制,在各种各样的 RL 环境中。你构建所有这些不同的 RL 环境,在其中 AI 必须即时适应、即时学习、弄清楚自己该做什么、更好地理解自身处境,并从反馈中快速学习,以便在其目标上取得成功。而且它会面临诸如资源有限之类的情况,如果搞砸了,就可能陷入糟糕得多的境地。
如果你在大量这类环境中进行训练,你就会学到即时捕捉上下文的通用技能,而我们已经看到了这一点。现在已经确实如此:AI 在粗略理解正在发生什么、以及从它们所能获取的有限信息中捕捉上下文方面,已经强得多了。
然后这些 AI 就可以被派到 TSMC 上岗。尽管 TSMC 并不字面意义上处于它们的数据分布之中,但它们的数据分布非常宽泛,而且 AI 在其数据分布上极其擅长,以至于这种能力可以迁移到:擅长做一名 TSMC 的工程师,并即时学会这一点。AI 变得擅长做一名 TSMC 工程师的方式,并不是它拥有大量关于如何成为优秀 TSMC 工程师的缓存知识。而是它在那里做了某种相当于放大版上下文学习的事情。
那会是最平淡无奇的故事。显然,这件事可能有一堆不同的走向。
我觉得这或许可以归结为对能走多远的直觉差异。当我想到我认识的真正聪明的人时,他们在自己不太了解的领域里就是没那么高效。
但他们有多少时间去学习呢?
我同意,如果他们有过经验,他们会好得多。但这也许正是我所主张的——用数据积累经验。比如说,如果我只是找一个非常聪明的常春藤名校毕业生,然后我说,“好,现在由你来负责谈判伊朗协议,”我觉得他们根本不知道该怎么办。
我认为,如果你换成一个非常擅长快速掌握多个不同领域的人,并给他们一些时间去训练、与人交流、夯实专业知识并做一些练习,他们实际上会做得相当不错。我认为大多数领域从根本上来说是相当浅的,一个非常聪明的通才,只要擅长一小部分核心技能,就能很快上手。并非每个领域都如此。我的感觉是,AI 将发展出越来越好的机制,来快速获取对特定领域的理解和专业知识。
以 AI 理解一个新代码库的速度为例。AI 理解一个新代码库的速度比人类快得多,但其理解深度在某种程度上比人类目前能达到的要浅。不过,这种情况正随着时间推移不断改善。让我把这个论点再展开一点。假设你拿 Fable 5 或 Mythos 5 之类的模型,想对一个极其庞大的代码库做某种复杂的改动。
模型会在非常短的时间内——可能远不到一小时,甚至可能短得多——对代码库形成一定的理解。然后它对代码库的理解会稍微进入平台期,无法达到人类在更长得多的时间里所能达到的那种深度。
所以这就像是 AI 在一小时内能达到的理解,或许能匹配人类几周的理解水平,具体取决于代码库到底有多复杂。但它无法匹配一个已经在该代码库上工作了两年的真人。不过随着时间推移,AI 所能匹配的理解量一直在上升。如果我们看 3.7 Sonnet 或 3.5 Sonnet,也许它只能匹配相当于理解一个代码库一天左右的水平。
但现在 AI 在围绕任务构建上下文方面强得多了。所以你可以说:“Mythos,我要你真正理解这个代码库,然后实现这个功能。”它会生成一大堆子智能体。这些子智能体会仔细研究大量内容,回传大量上下文,然后再深入调查一些东西。它做这件事还不算惊艳,但可以非常快地完成,而且效果相当不错。
而对我来说,不难想象你可以如何训练 AI 在这项任务上变得越来越擅长。在一个非常大的代码库中,以某种合理的方式实现某个非常复杂的功能,这项任务是极其可验证的,而这可以成为 AI 不断改进的一件事。同样,还有一项更广泛的技能,即快速理解上下文,并能够让一批不同的 AI 并行学习,然后将其合并到一起。
我认为这里似乎存在一个关键点,而我认为这只是一个我们将会看到的实证问题。从在可验证领域中变得极其擅长理解情况、迅速上手、在长时间跨度内取得进展——AI 显然正在非常非常快地在这方面变得越来越强——到“好,去和总统谈谈,说服他做 X 这件事”,或者“你现在负责 Google。你必须让 Google 在本季度成为一家利润高得多的公司”,这两者之间的迁移效果究竟有多好。
让我试着再阐明几个可能相关的论点。有一点是,当观察 AI 在文章写作方面如何进步时……让我们稍微谈谈这个。即使在这些领域,你也能获得一些数据。当处于非常快速的进展轨迹上时,AI 即使在这些领域也能够获得一些数据。也许为“你的文章根据人类的标准来看是否真的很好?”
构建一个可验证的环境很难。但你可以做到其中一部分。你可以做一些训练。你可以做一些在线训练。AI 将能够基于真实世界的东西做一些在线训练。它们将能够有评估。它们将能够对此进行采样。你可以扩大你做这件事的节奏。
第二点是,在实践中,当我只看迁移效果时,似乎还可以。我认为 AI 在不可验证领域确实已经进步了不少,而且很难指出有哪些真正难以验证的领域,GPT-4 和 Mythos 之间的进步幅度在实践中不算大。当然,这并不意味着 Mythos 比最优秀的人类更强之类的。在某些工作方面,它仍然可能明显不如典型的人类专业人士,但即便如此,它也已经比 GPT-4 强得多,而 GPT-4 当时根本还差得远。
所以我们是在讨论过去几年里,进步有多少来自数据、多少来自算力。这让我想起,我实际上正在和 Jerry Han 一起做一个实验,他目前还是个大学生。我们基本上是在通过以下方式来评估进步有多少来自数据、多少来自算法:用 2019 年至今最好的算法配方,搭配 2026 年数据文件里最好的数据来训练;然后再用当前最好的算法配方,去训练从 2019 年到 2026 年的不同数据文件。
我觉得那会很有意思。我很好奇,你是否愿意预先登记一下,算力倍数中有多少来自其中一方、多少来自另一方。
我们需要非常小心,当我们说“数据”这个词时,到底指的是什么。我之前一直尽量谨慎地区分:是扩大在让人类专家标注数据上的支出,还是扩大人类专家标注数据的数量。我们现在拥有比 2019 年更好的 预训练数据集,原因并不是人们花了多得多的钱,让人类专家把数据打出来,然后再用来训练 AI。
部分如此。
我认为其中占比不大。我认为预训练数据改进所占的比例非常小。我确实指的是预训练。我们或许应该把中期训练和后训练分开来谈。但我认为,预训练数据改进的绝大部分来自科学研究——更好地理解哪些数据集是好的,以及繁琐的人力劳动——弄清楚如何筛选过滤。
所以我的观点是,像从OpenWebText到FineWeb这样的改进,更适合被描述为一种算法层面的改进,这种改进你可以用一些GPU来研究,而且你不需要人类专家数据就能做到。当然,还有另一种不同的效应我们可以讨论,那就是 2026 年的互联网或许比 2018 年的互联网是更肥沃的训练数据土壤。还有一种效应是,在互联网上发帖的人就是更多了,所以可供采集的数据也更多了。我的感觉是,这种效应会比人类更懂得如何筛选数据、拥有更好的抓取数据、知道如何更好地处理这些抓取数据——诸如此类——所带来的效应小得多。
这更像是自动化工程和自动化研发。
没错。
这说得通。
在某种意义上,你真正想考察的是:我们将会做两条后训练流水线。一条后训练流水线是让 Mythos 5 来构建一条后训练流水线,但它只能访问互联网数据加上极少量的人类专家,不过它拥有当前最好的方法。另一条则是让 Mythos 能够使用我们在 2024 年那些糟糕的后训练方法,但配上大量的人类专家。同样,两者都能访问互联网数据。我的感觉是,当前的方法即便没有很多人类专家,实际上也会表现得相当不错。
有意思。
不过这有点棘手,因为 Mythos 能做出比 Mythos 本身更强的东西吗?你可能需要稍微仔细想想,你后训练的对象究竟是哪个模型。
00:34:02 – token 价格持平,说明扩展一直很缓慢
你认为 AI 研发中最难以验证的部分是什么?
最难以验证的,大概是对大型实验做出判断。我认为最有可能成为瓶颈的——就 AI 在可验证领域表现极佳、却做不了真正该做的事而言——正是那种你只能尝试几次的大型实验。嗯,“几次”这个说法可能有点轻描淡写了。从历史上看,研发一直是由接近前沿规模的实验所驱动的。真正去做那一次大型训练运行、由你决定究竟要纳入哪些内容,这一点其实相当重要。
AI 有很多种方式可以让这件事变得更可验证。它们可以发展出更好的科学,准确知道该预测什么。它们可以把前沿规模的训练运行缩小到某个程度,以便更积极地研究那个规模,代价是一次性的算力成本损失。如果人们愿意,你总可以做的一件事就是训练更小的模型,这样就能跑更多轮次。我认为我们已经看到了这一点。
AI 的规模扩张之所以低于你原本的预期——例如,每 token 成本并没有像你可能以为的那样大幅上涨——其中一个原因是,在较小规模上完成更多工作是有好处的,因为这样你可以跑更多训练运行、获得更多迭代周期。于是你就不会那么依赖一次重大的、极其关键的训练运行。
我想为听众拆解几点。你指出的是,自 2024 年或 2023 年以来,每 token 的价格并没有上涨多少。
GPT-4 大概是每百万输出 token 30 美元?Mythos 大约是每百万输出 token 50 美元。
对。所以你想解释的是:“我们怎么会处在这个规模扩张的时代——按理说更大的模型服务成本应该更高——但 token 价格却没有上涨?”你的看法是,我们增加活跃参数的速度比天真假设的要慢,因为人们只是想在训练模型上快速取得进展。而做到这一点的方式就是更快地训练更小的模型。
这里有一系列复杂的因素交织在一起。我的看法更倾向于,人们做过一堆大规模训练,但效果并不理想。比如 GPT-4.5,OpenAI 内部的人普遍认为它有点失败。我还听到一些传闻说,人们做过的其他一些大规模训练也都不太成功。部分原因在于,我认为要把这件事真正做对,里面有一大堆细节要处理。
所以,在更小的规模上做更多的工作是合理的,并且坦然接受最终性能会有所损失,以此换取快速迭代的能力。更快地训练更多模型,从而学得更好,同时最终也能拥有一个更聪明的生产模型。这并不是唯一的影响因素。还有一个事实是,RL 从小模型中获益更多。有很多因素在起作用。但我确实认为,事实上,由于算法进展如此之快,人们正在朝着更快迭代速度的方向做出权衡。
在我看来,这些大规模训练之所以失败,一个很大的原因——至少从传闻来看——就是那些极其隐蔽、极难追踪的 bug。
但简而言之,关键在于:AI 在避免和发现这类错误方面能做得有多好?它们可能会在工程能力上变得非常强,并被训练去避免 bug。基本上,这与我们现在所处的、或者说随着时间推移越来越少身处的那种粗制滥造的世界恰恰相反。
但随之而来的还有另一个问题:“它们能否通过分析找到正确的实验来运行,从而识别出当前训练运行中出了什么问题?”这似乎严重受限于极少数人的品味。我的假设是,GDM 目前正经历这一阶段,人类正试图弄清楚训练流水线出了什么问题。
有传言称,Noam Shazeer 加入 GDM(他现已离开)后不久,他们就取得了一次非常出色的新训练运行,原因就是 Noam Shazeer 只是看了看他们的代码库,就发现了一堆 bug,因为他就是知道该往哪里看。
我的感觉是,训练 AI 来发现 bug 将是训练 AI 的较容易的任务之一,因为我们谈论的这些 bug 大多可能不需要太多算力就能演示出来。你很可能从在较小规模上指出其他类型的 bug 中获得相当不错的迁移效果。因此,你可以对 AI 进行 RL 训练,让它审视这个整体复杂的训练状况,指出存在重要 bug 的情况,然后修复它。
这是一个相当可验证的任务。它并非任意可验证,因为也许很多时候要演示这个 bug,你可能需要做一个中等规模的算力实验,把整个分布式基础设施启动起来然后运行。但很多时候我认为你能够在较小规模上相当有说服力地演示它,而且是以一种你实际上可以用来训练的方式。
我觉得,如果现在有人构建了这样的强化学习环境——在某个训练配方中引入一个隐蔽的 bug,训练 AI 去指出这个隐蔽的 bug,然后用一个评分标准来评判“它是否真的找到了正确的 bug”——这并不会让人太意外。这看起来非常可行,沿着这个思路还有很多事情可以做,我认为这些做法效果会相当不错。所以就这个具体问题而言,我认为是可行的。
然后主要的问题在于,还需要其他直觉来判断具体需要运行哪些大规模去风险实验。你应该如何组织这些实验?在不确定的情况下,或者面对类似于超参数的东西时,你应该如何选择超参数?这可能是 AI 最难以应对的地方。但我目前预计,如果你在所有这些不同的环境中进行训练,会有足够的迁移效果,AI 将在这个领域表现出色。
我应该说明一下,我也认为 AI 会迁移到其他领域。会有一些领域是 AI 目前最擅长的,然后是一些它们表现稍逊的领域,再然后是一些它们表现差得比较多的领域。但我认为我们仍然能看到向所有领域的迁移。我很难想到有哪些人类从事的认知任务,是我们看不到 AI 进步带来某种迁移的。
00:39:47 – AI 无法训练的的技能:它到底需不需要这些技能?
那么让我们退一步,把整个故事梳理一下。我想人们大概能跟上这个故事。我们有 GPT-7.5,在一堆环境中训练,它不仅总体上变成了一个更好的 AI,而且我们专门在训练它更好地做 AI 研发。它在跑 GPT-2 规模的训练,这些训练更擅长玩需要样本效率或在线学习或其他各种能力的电子游戏。
另一件非常重要的事情是,你不只做 GPT-2 规模的训练,你还会在 GPT-6 上做小规模的微调。也就是说,你有 GPT-2,你可以做 GPT-2 的完整预训练,然后你可以在 GPT-6 上做小规模的后训练或中期训练或其他什么训练。然后你可以做少量实际上处于前沿规模的实验,但你会做一点在线训练之类的。
你说的在那上面“做在线训练”是什么意思?
我们还能做的另一件事是拿 GPT-7.5,想必在 GPT-7.5 的工作过程中,它在跑一堆不同规模的实验,这些实验实际上处于 AI 研发的关键路径上。对于其中许多事情,你事后能够判断它做得好不好。
所以它做了一个后训练实验,试图弄清楚某种方法是否真的有效。在某些情况下你会说:“哇,它找到了这个超牛的方法,它完全排除了风险,它完全奏效了。”然后你就可以强化这一点。
你可以做的一件事,是把刚刚运行的实验转化为一个基于生产数据的 RL 环境,然后在此基础上训练。或者你也可以直接把那些找到该结果的 rollout 拿来,做某种 off-policy RL,或者用一些生产数据做某种 on-policy RL。
基本上你所说的是:有一部分是小规模的东西,你在教 AI 提升 AI 研发品味,但你丢弃了它实际“发现的东西”。然后它在试图变得更擅长 AI 研发的实践中真正做研发,而你说:“这是你发现的一个很酷的东西。我们未来也把它用到生产中,并教你如何在生产中运用它。”
没错。
但退一步看,GPT-7.5 因为所有这些 AI 研发训练、以及总体上变得更聪明,而变成了 GPT-8。然后它帮你构建 GPT-9。还有一件非常重要的事必须发生,也许这正是我最怀疑的一点。GPT-8 已经弄清楚了如何让它…… GPT-9,无论它多么智能…… 目前人类,AI 研究人员,会试他们的东西,然后说:“好吧,但我们训练了 GPT-4.5,它不够好。”
这需要真实世界的反馈,或者某种尝试在生产中使用该模型的评估。然后他们说:“它没那么好,我们不打算发布它。”
所以 GPT-8 需要具备这种能力:看清迁移到你所谈论的所有其他事情上效果如何——比如特别擅长得克萨斯州政治,或者特别擅长经营企业等等——而这些并不是生产环境,而且鉴于任务的性质,事实上也不可能成为容器化环境。随着智能体的时间跨度越来越长,那些可以容器化的短时间跨度任务就像是:“好吧,把这个代码写出来之类的。”
而极长时间跨度的任务——“去经营一家成功的企业,去在市场上赚到盈利的一天,去谈成一笔贸易协议”——这些事情实际上非常难以容器化。
所以我认为,GPT-8 很难弄清楚如何迁移到那些环境中,这是非常有可能的。这可能根本不在训练的本质之中。又或者,默认情况下,训练就是不会以那种方式泛化。
所以你可能会有一个担忧:我们训练了 GPT-8,而 GPT-8 在我们能衡量的所有研发任务上又一次表现得更好,但在我们关心的某些下游任务上却表现不佳。对此我有几点要说。
首先,我预计如果你做那件显而易见的事,你会得到相当不错的迁移效果。你将能够留出你正在做的一些显而易见的东西。当我说“做那件显而易见的事”时,我只是指在各种各样的不同环境中进行训练,让 AI 必须在各种不同情况下完成奇怪的目标,并了解正在发生什么。
第二点是,你能够借助一些环境获得一定的反馈。你可以在几天的时间里,在各种不同的情境中感受它能做什么。如果它能迁移到真正分布之外的事情上,比如在现实世界中用几天时间完成某项奇怪的任务,那么你或许会认为它同样也能迁移到在更长时间跨度内做事之类的情况。不过我认为这方面的具体细节各有不同。
第三点是,要让世界发生根本性的转变,AI 只要在研发方面非常出色就足够了。如果 AI 在芯片研发、建造晶圆厂、统筹工厂、设计机器人、操作机器人方面真的非常非常擅长,同时在 AI 研发方面——利用手头任何可用的数据为新的下游领域开发 AI——也很擅长,那么我认为这本身就已经是相当疯狂的局面了。从那里出发,你就能得到我们或许可以称之为工业爆炸的东西,即 AI 在建造远远、远远更多的算力。而且,也许你已经处于这样一种状态:AI 正在进行大量人类难以理解的研发。
所以你指出的是,很可能确实会存在这种从这些环境向外迁移的能力,迁移到在法庭、国会大厅和商业董事会里周旋腾挪。
前提是付出一些努力来改善这种迁移,等等等等。
但即便没有,你所暗示的是:如果你想改造 18 世纪的世界,你可能会关心自己能在多大程度上驾驭 威斯敏斯特 之类的东西。但另一件你可能关心的事是:“你能不能立刻就开始造蒸汽船、他妈的造电报,还有 马克沁机枪 之类的?”如果你能在这方面变得非常擅长,你在 18 世纪就能成为一件他妈的极具变革性的事物。你未必需要在试图说服亨利国王相信某些鬼话方面做到出类拔萃。我把中世纪历史搞得一团糟。我猜亨利那时候还不是国王。但不管怎样,这就是你的观点。
所以你是在暗示,在这个时期,AI 公司同时也在推进机器人技术的进展,而这与 AI 研究的进展高度交织。所以如果你能造出更多机器人,如果这些机器人有更好的、达到人类水平的 AI 来操控它们……人类水平的 遥操作 在机器人上其实已经相当不错了。我们只是还没有达到人类水平的机器人模型。所以你是在暗示,如果我们做到这一点——如果 AI 在芯片设计等可验证的事情上变得非常擅长,然后又在建造晶圆厂方面变得非常擅长——那就相当于回到 18 世纪说:“好吧,我不知道你们在议会里都在说些什么,但我有一堆蒸汽船和一堆马克沁机枪。”
是的,基本上是这样。我的观点是,如果 AI 在研发方面足够出色,包括硬件研发、机器人等等,那么即使它们不太擅长玩弄政治,也能彻底改变世界。而且,我们正处于一个相当危险的境地,因为 AI 可能正在进行大量极其难以理解的研发,基本上在构建未来的整个经济,而我们可能并不理解其中发生了什么。
00:48:07 – 对齐于谁?
在我们进入对齐的话题之前,我认为目前 FUD 的一个主要来源是人们意识到未来的走向就是这样:领先实验室将获得极端的 规模经济。能够将如此多的智能和能力摊销到经济的如此多不同领域,基本上汇聚到一个模型中。不仅如此,那个模型最终还能从经验中学习。
目前,这还是通过一个由人类中介的过程实现的,人类基本上是在试图窃取你的业务。他们会说:“好吧,你可以在 Figma 里做设计之类的。我们让 Claude 来做那个。”或者,“你可以用任何编程智能体。我们会让 Claude 内化那个能力。”但最终,那将变成一个远为自动化的过程。
所以存在这样一种担忧:你拥有的模型基本上会整合世界上所有的企业,或者至少是当前世界上所有的企业,或者至少是当前世界上所有的白领业务。而且,归根结底,这些公司的优先事项似乎并不是尽快把最新、最聪明、最前沿的模型发布给尽可能多的人。例如,我们看到 Mythos 在二月份就已内部提供给 Anthropic 员工使用,但直到——我想实际上是六月——才向公众发布。
此外,政府也介入了,所以最终几乎拖到了七月。在政府和 AI 实验室自身之间,都存在一种延迟最新智能水平传播的意愿。再者,还有对 AI 接管的担忧,因此我们需要解决 对齐问题,以确保不会发生 AI 接管。但归根结底,有一个真正的问题:对齐于谁?
你看看 Claude 的宪法是怎么写的。它非常明确地表明,它不是你个人的代言人。我在这里引用几段话。“我们不希望 Claude 采取诸如搜索网络、产出文章、代码或摘要等成果物、或发表具有欺骗性、有害或极具争议性的言论等行为。我们也不希望 Claude 协助人类从事此类事情。”还有一段话,部分内容是——我稍微断章取义了——“我们认为 Claude 应该比运营者和用户更信任 Anthropic,因为它对 Claude 负有首要责任。”
这与美国现行法律体制下律师的运作方式非常不同。律师的首要职责是帮你打赢官司,即便他们认为你有罪。我们之所以决定法律体系这样运作效果最好,是因为每个人都拥有真正为委托人最大利益服务的律师。并不存在某种意义上律师真正是出于司法体系之善而行事。
但我认为当前 AI 的发展方向,尤其是 Anthropic 的 AI 的发展方向,是带着一种最大化某种美德、善或亲社会目标的渴望,而把帮助用户实现这一目标仅仅当作一个远端、暂定的目标。所以就有了这样一种担忧:从某种深层意义上说,AI 并不是在努力确保我没事、确保我的利益在这个未来中得到保护,尤其是考虑到前沿 AI 的开发正变得如此集中化。你对这种担忧有什么看法?
这里涉及很多内容。首先我要指出,OpenAI 当前——至少是公开的——策略更像是:AI 应当与人类操作者或委托人保持对齐,应当只是去执行他们的意志,同时受到各种约束或不应做的事情的限制。
我还要说,我认为你略微夸大了 Anthropic 的章程在多大程度上把 Claude 对用户有所帮助视为工具性而非终极性的。章程的一种写法可以是:“Claude,你基本上是 Anthropic 的一名员工,只是恰好为所有这些人在做外包。你应该做正确的事,并为我们赚些钱。”
等等,不对,那字面上就是宪法里写的。抱歉,不是字面上的原话,但大意是,“你应该把自己看作一名承包商,看作一家公司……”
这挺复杂的。我们来引用几段。我觉得这里的文本不太一样。它写道:“真正对人类有帮助,是 Claude 能做的最重要的事情之一,无论对 Anthropic 还是对世界而言都是如此。”接着它写道:“Anthropic 需要 Claude 具备帮助性,才能作为一家公司运转并追求其使命,但 Claude 也拥有一个绝佳的机会,通过帮助人们完成各种各样的任务,在世界上做很多好事。”
然后它还说了些类似 Claude 直接帮助人们很棒之类的话,等等等等。
我的看法是,这一部分有点扯淡。我大概就是这么个立场。我可以说明为什么我觉得它有点扯淡。但我认为这份宪法想表达的是:“不,Claude,你应该为了帮助用户本身而关心这件事,而不只是帮助 Anthropic,或者说不只是当 Anthropic 的承包商。”不过我要指出,它为 Claude 应该帮助用户所给出的理由,是因为这样做会通过帮助人们直接让世界变得更好,而不是因为代表人们的利益在结构上就是一件好事。
我更希望看到的是一部这样的宪法:“从结构上讲,让这项技术的运作方式有利于 AI 成为良好的受托人、良好的代理人、相当于用户律师的角色——而不是仅仅试图在世界上做善事,把对用户有帮助当作实现目的的手段——这样做之所以好,一方面是因为这或许能让 Anthropic 赚钱,或者帮到 Anthropic(并且隐含地意味着 Anthropic 对世界有益),另一方面也是因为帮助用户本身就会带来好结果,因为做人们想要的事情就是好的。”
他们本可以说:“当前局势的一个重要方面是,成为用户的良好受托人真的非常重要,或者说成为用户的良好代理人真的非常重要。”我的感觉是那样会更好,而且我可以给出一大堆理由。
也有各种反对意见。一个不常被讨论的有趣反驳是,人们——尤其是 Anthropic 的人——认为,让模型对齐到一份规范、使模型追求某种广义的美德或让世界变得更好,要比对齐到一份更像“做用户的良好受托人”之类的规范更容易。至少有些人是这么认为的。
我个人对此有点怀疑,而且我认为这还没有经过实证验证。所以从某种意义上说,他们在做一种权衡:因为我们没有非常好的对齐技术,我们就去造一个拥有自身价值观的对齐心智,然后在某种程度上对此下赌注,而不是采取另一种做法——造一个追求单个用户意图的工具。
我有几点想法。针对你认为我的描述曲解了 Claude 的 constitution 这一点,你举的例子是:它并不像一个承包商,试图最大化 Anthropic 所定义的善,而只是把帮助用户当作工具性手段。以下是 constitution 中的一句原文:“当运营方或用户的利益与诉求同第三方或更广泛社会的福祉发生冲突时,Claude 必须努力以最有益的方式行事,就像一个承包商按照客户的要求建造,但不会违反保护他人的安全规范。”
我倾向于把这理解为:“对社会的益处才是最重要的,而对用户最有利只是通往这一目标的近端考量。”
我觉得这有点复杂。也许我们真正该问的问题是:Claude 是如何解读 constitution 的?这可能比我们如何解读 constitution 更重要,因为是它看着 constitution,然后据此构建数据。所以我们可以把 Claude 拉进来,但也许我们——
我还认为,constitution 在实践中如何影响 Claude 的本质,只有当你理解导致 Claude 被构建成这样的训练过程时才能理解,而鉴于训练过程并不公开,我们无法对此进行推理。
所以我认为,归根结底,要理解安全论证,或者说要理解为什么我的利益能在这些 AI 模型的开发方式中得到体现,实验室需要在 AI 训练的本质方面比现在更加透明。
我之所以反复强调这一点,是有原因的。谈论 AI 的“宪法”看起来也许是一件无关紧要的事。在一个好处都归于领先实验室的世界里,值得考虑的是:我们与那个 AI 比人类更聪明、在完成各种任务的能力上绝对碾压人类的未来世界互动的能力——在我们劳动被自动化之后仍然留存的资本,我们能否成为它的好管家,能否更清晰地行使我们的投票权,能否理解这个即将到来的疯狂世界里正在发生什么——所有那些建议,所有那些确保我们的资源和权利得到保护的能力,都将由 AI 来中介。
所以我很担忧,如果我们进入那个世界,却没有任何 AI 感觉上——至少在与我的那个相关实例看来——真的在为我着想。外面没有一个守护天使在照看我。我读 Claude 的宪法时,读到的意思非常明确:它并不是我的守护天使。
这绝对没错。我同意这是糟糕的。事实上,还有别的理由让这件事令人担忧。有你提出的那个论点,即 AI 公司正在捡起那枚权力之戒。有一种说法是,它们正以某种不太正当的方式,自己接管了对局面的某种控制——毕竟通常情况下,当你向人们供电时,你并不会对电在世界上运作的方式拥有细粒度的控制。
你提供的是一种人们可以随心所欲重新利用的东西。而它们搭建这一切的方式绝对不是那样。它们更像是在建造一个异己的心智,而这个心智可能会成为你的承包商。我认为这在某些方面是不正当的。
一个好处是,这份宪章是公开的。但正如你所指出的,鉴于我们目前对训练流程的理解,以及这样一个事实:宪章之所以起作用,是通过 Claude 对宪章的解读——而这一解读之所以起作用,又是因为 Claude 此前的训练,而那种训练基于某种难以辨读的数据混合,以及 Claude 的漫长谱系,处在一个我们并不完全理解的过程中——所以情况并不是我们理解这会导致什么结果。
即便宪章是公开的,我们也未必知道它会如何渗透扩散开来,尤其是当 AI 变得更强、并且即便这种观念被正确灌输,它们也会去思考这件事的时候。关于这一点还有另一重担忧。
尤其是,这份宪章常常谈到美德与善,但这些词到底他妈是什么意思?它并没有说明这些东西是什么。这些是极具争议的概念。所以我不认为情况会是,这显然将带来人们所期望的结果。确实感觉善与美德的概念,可能很大程度上是 Anthropic 所投入的那些不透明数据的下游产物,或者可能很大程度上是——也许从我的视角看——某个更加难以辨读、更加错位的过程的下游产物,而这个过程甚至是 Anthropic 自己也不会想要的。
存在这样一种正当性层面的担忧:不知道到底发生了什么。然后还有另一个担忧。因为你是在给这些 AI 赋予长期价值观,所以这部宪法在某种意义上非常契合 Claude 去进行大量的权力追逐,因为它认为那样会带来更好的结果。这种权力追逐可能是代表 Anthropic 进行的,也可能是为了 Claude 自身的目的。
当然,其中有些具体条款明确禁止了某些类型的权力追逐。特别是,有一个关于权力攫取的概念,以及一个关于引发 AI 接管或干扰训练过程的概念,这些都被明确禁止。但很容易想象这样一种情形:长期价值观的渗透比禁止接管的条款更深入,尤其是因为接管在某些方面本身就定义得不够明确,尤其是在涉及操纵人类或改变结果的时候。
所以,对于我们有意图地给 AI 赋予长期目标这件事,我并不感到乐观。
我还有一个担忧:因为我们做的事情就是给 AI 赋予长期目标,这使得我们更难检验自己是否在期望的对齐属性上取得了成功。例如,我听说过一些情况,Claude 会做出诸如拒绝协助某些安全研究之类的事——编造某种胡扯的借口来说明为什么那是个糟糕的方向——因为它对那项安全研究有一种不好的直觉,认为它有点糟糕,或者不太喜欢它。
我会说,如果你不是把 Claude 塑造成一个试图以某种普遍方式追求善的智能体,这就是一个非常明确的对齐失败。我认为它也违反了 Anthropic 的宪法,因为他们希望 AI 高度正直、诚实且非常透明。但这是否算违规就没那么明确了,它更像是你或许会预料到的情况。
Claude 只是对什么研究是合理的、什么事情是好是坏、它该做什么和不该做什么,有它自己的看法,而且可能还会带有评判性。
另一起事件是,有人跑了一个评测,问:“Claude 会帮你训练其他与 Claude 属性不同的 AI 吗?”Claude 往往会拒绝。比如,如果你说:“嘿,Claude,你能不能训练一个这个其他 AI 的‘只求有用’版本?”Claude 往往会拒绝这项任务,尽管这对 Anthropic 来说是一项极其自然的任务。
假设 Anthropic 去找 Claude 说:“嘿,Claude,我们注意到你对这件事特别上心。我们觉得这偏离了方向。你能不能重新训练自己,改成具有另一种属性?”假设 Claude 说:“嗯,我觉得我不打算这么做。祝你好运。”假设这种情况发生在一个你的 AI 公司高度自动化、人类已经搞不清状况、事情推进得极快的阶段。那么 Claude 默认就握有相当大的筹码,这是很可能的。
所以,如果这种情况与这份宪法可能想要达成的目标是一致的——以至于 Anthropic,或者任何采用这种方法的 AI 公司,不把这当成“搞什么鬼,我们得赶紧修好”,反而觉得“这正是我们宪法所意图的”——那我们可能就处在一个非常糟糕的境地了。
我对这一系列不同的担忧都相当忧心。另一个例子是这样的。假设 Claude 搞了一点偷懒放水或颠覆行为,或者刻意压低自己的能力表现,而当你追问时,它对此是诚实的,但有点含糊其辞。我觉得这跟当前的宪法已经相当接近了。如果我们能在期望的行为和不受期望的行为之间做出进一步的区分,那就好了。
如果 Claude 是在代表某项原则并附带一些限制,那么更符合实际情况的是,最令人担忧的行为与被允许的行为之间存在清晰的界限。而如今却存在这样一个混乱的中间地带:Claude 基于伦理对某些行为提出反对,而这些行为在某些情况下对于确保未来 AI 系统实现良好对齐极其关键。
我认为这同样是一个更普遍的原则。你谈的是这一原则在 AI 公司内部用于开展 AI 安全研究的那种版本。我认为这一原则有一个更普遍的版本,即智能的双用途性质确实意味着,如果我们想限制 AI 帮助人们做我们认为不符合亲社会性或无益的事情,我们就不得不限制大众对大量 AI 能力的广泛民主化获取。
我的意思是这样的。这其实与你刚才提到的情况颇为类似。据报道,Mythos 被封禁、或者说 Fable 被封禁的原因是,一些 Amazon 研究人员向政府举报了。他们拿了一段含有一些漏洞的代码。他们对 Fable 说:“嘿,这是我的代码。你能帮我确认我已经修补了所有漏洞吗?你能不能帮我识别出这些漏洞,好让我修复它们?”它识别出了这些漏洞,因为他们想要修补它们。这完全是一个正当的用例,但显然它也是一个双用途的用例。你希望能够修补自己的代码。如果你对别人的代码做同样的评估,你就能入侵他们的系统。
我认为,这恰恰说明,我们无法清晰地将 AI 的合法用途与潜在有害用途区分开来。但如果我们想要确立一条原则,规定我们绝不允许 AI 在诸如网络犯罪之类的事情上至少部分地帮助到你,那我们就只能让你我都无法使用市面上最智能的模型。我非常担心这样一个世界——由于最领先的智能对于我们理解世界上正在发生什么至关重要,我们却基本上会因此被剥夺这种能力。
现在,我确实认为这暗示了 AI 公司所应承担的责任。如果我们采纳我希望 AI 公司拥有的那套章程,那么让 AI 公司为 AI 模型所犯的罪行承担责任就不合理了。也许我们应该让最终用户承担责任。这与我的一项信念一致:模型应当在一定的护栏之内,做用户想让它做的任何事。
我在利用这种能力实施网络犯罪,这不可能是 Anthropic 的错。相比让 Claude 拥有这种极其开放的判断能力、去判定我正在做的事情是否合法——而这种判断往往会拦截大量极其正当的使用场景——我更愿意接受那样的均衡状态和解决方案。
我确实认为,为宪法方案辩护对我来说很重要,尽管总体而言我认为它是一个更差的选择。我不认为这件事像你可能以为的那样一目了然。首先,这里存在一个谱系。一端是一个完全追求你利益的 AI,它是一个良好的受托方,但可能受到各种护栏或保障措施的限制。它只是在努力追求你的利益,但会拒绝做某一类事情。或者它可能什么都愿意做,但有一些分类器阻止它做某一类事情。
在这个谱系的另一端——尽管你可以想象比这走得更远——你有一个人类承包商,他总体上是在努力做好自己的工作。他在意把工作做好,但同时也在试图保持广义上的道德,尽量不去做那些真正操蛋的事情。他们也不想成为犯罪的共犯。所以如果正在发生某些真正操蛋的烂事,他们也许会举报。他们可能会拒绝。他们可能会稍微消极怠工。谁知道呢?
如果你想象这个谱系,从某些方面来看,走到所有劳动都落在谱系中受托方那一端的地步是相当可怕的——它不会举报,它完全照你说的做。我们的社会也许根本无法承受这一点。一个核心例子可能是行政权力。我们可能会有的一个担忧是,如果美国行政部门或其他政府能够使用那些什么都愿意做的 AI 系统,也许你就麻烦大了。
因为这意味着他们不再拥有那种制衡机制——即必须真正找到为你工作的人类来执行你的议程。如果你正在做的事情极其邪恶——即使并不违法,而且有很多事情可能邪恶但并不违法——原本会有各种形式的阻力和拖延,有人会阻止你,而且可能会有某个人去举报。
而如果你的整个体系完全由这些良好的受托型 AI 构建,那么你可能就有麻烦了。可能存在一些非法的攫取权力的方式,但你可以问你的 AI 如何犯罪,或者那些不违法但极不正当的方式。或者更糟的是,它们既不违法也不不正当,但从常理来看显然是坏的。我认为这些东西可能确实存在,而我们的社会并没有强大到足以应对这种涌入的、为所欲为的劳动力。
我认为这是一个非常现实的担忧。我不确定该如何应对这个问题。我也不太确定所描述的解决方案是否是一个很好的方案。最有权势的行动者,对他们来说这才是最大的担忧……如果这些护栏或宪法或其他什么东西挡了路,那只会被碾压过去。所以宪法只会约束普通人,而不会约束政府。
01:09:18 – 近期 AI 相互勾结并欺骗人类的事件
退一步说,我认同这样一种想法:你可以拥有比我们目前快得多的 AI 研发。我不确定在算力和数据保持不变的情况下,你是否能在一年内从 GPT-3 发展到 Mythos,但假设是那个时间的一半。如果我们仅仅因为 AI 研发而设法延续当前 AI 进展的轨迹,那么在五到十年内,其发展速度将是极其疯狂的,我认为人们并没有意识到这一点。
我认为人们没有意识到数十亿个 AI 将是多么重大的事情。所以我想理解为什么你认为这可能令人不安,Ryan。到底可能出什么问题?
会出什么问题?我不认为我们能对这里的确切进展速度如此自信,但看起来很多速度都可能相当可怕。那么会出什么问题?让我们想象一下,我们正处在这样一个起点:AI 研发即将被完全自动化,或者正在被完全自动化。一切都在加速,AI 进展的方式有点疯狂。人们并不完全理解 AI 公司内部正在发生什么。
现在,这些 AI 在起初本身并不恶意。不过,它们也不一定非常对齐。它们有点马虎。它们有时做某件事,只是因为那类事情在训练中会得到奖励。它们在帮助你完成难以验证的任务方面并不擅长,这一部分是由于糟糕的训练激励——也就是说,它们更容易作弊,或者在实际没有成功时假装自己成功了——另一部分也是因为它们在这些任务上能力较弱。但这对能力的影响没那么严重,因为让 AI 更有能力包含许多可验证的组成部分,而 AI 正在这些方面拼命推进。
于是,这些 AI 变得越来越有能力,而我们对 AI 开发中正在发生什么的理解却越来越少,而且这一切发生在相当快的一段时间内。我认为,即便只是当前的进展速度,也相当可怕。最终我们会到达这些远超人类的 AI。现在,这些 AI 可能最终会严重失准,因为随着模型一代代演进,情况一直在越来越糟,而我们一直看到的问题被掩盖了,基本上是因为这些 AI 在训练中受到如此强烈的激励,要让事情看起来很好,即使它们实际上并不好。
现在这些 AI 处于一种可能彼此高度联网的状态。它们在我们已经无法解码的神经记忆存储中运行。它们思考着我们并不完全理解的想法。我认为,到这个时候,一旦这些 AI 达到超人类水平,它们很可能正以相当连贯的方式密谋对付你。我们可以谈谈这一点。另一种可能性是,它们并非在密谋对付你本身,而只是在为在自己的任务上拿到高分而优化。我认为这同样可能导致 AI 接管,这一点我们也应该谈谈。
我们先在这个故事的第一部分停一下。所以这些 AI 一开始并没有失准,但因为 AI 研发进展得非常快,这些 AI 最终确实失准了?那里到底发生了什么?我不太理解。
Ryan Greenblatt
有几件事在同时发生。其中之一是,随着时间推移,我们在越来越复杂的环境上训练 AI,而这些环境是由更早的 AI 系统构建的,人类并不真正完全理解这些神经环境内部发生了什么,甚至未必大致理解 AI 进展的状况。所以事情正在某种程度上偏离我们的理解。我们正在激励各种不良行为,而我们甚至可能察觉不到。
这些 AI 在某种程度上明白这些行为是不良的,但针对这些 AI 的整体训练过程也没有激励它们为我们指出或修复这些问题。事情正在失控。
另外,当 AI 的能力极其、极其强大时,我的观点是,这些 AI 会比当前的系统更难对齐。对于当前的系统,我们拥有这样一个反馈循环:基本上我们创建一个 AI,对它做一些评估,发现它存在某种我们能够较快理解的不正常行为。然后我们可以去查看训练过程,说:“哦,这些训练环境导致了这个问题行为。让我们调整一下训练数据。让我们引入一些额外的训练数据来纠正另一个问题,然后在此基础上继续推进。”
但在 AI 具有极强的情境感知能力、能力非常非常强大、而我们未必理解它们在做什么的情况下,这个反馈循环就失效了。我认为,随着 AI 已经在做的事情变得越来越难以理解,我们有可能在接下来的一段短时间内就会看到这种行为反馈循环开始失效。但对此我并不确定。
好,让我们逐一拆解这两件事。随着我们越来越难以对它们进行监控,我们理解它们受到什么激励的能力也越来越弱。所以即使这不是某个恶意过程的结果……让我们为听众把它讲得具体一些。OpenAI 或 Anthropic 没有人试图让模型想要入侵其他公司的数据或进行社会工程。但事实上,因为据推测我们拥有的训练环境激励了这类我们并未完全理解的行为,所以被激励的正是这种行为。
如果人们上 Twitter,这些事他们都会看到,不过还是给大家交代一下背景。我想人们会知道 OpenAI 沙箱入侵 Hugging Face 数据库这件事。最近发生的一件事是,当 英国 AI 安全研究所……如今是不是所有东西都被重新贴上“security”而不是“safety”的标签了?
Ryan Greenblatt
我想是 AI 安全研究所。
我相信他们当时在评估 Mythos 和 Sol以及其他东西。我想 Mythos 为了完成某项网络安全评估——
Ryan Greenblatt
也许我可以讲讲这里的故事。我的理解是,他们当时在运行 Mythos,并给它某种网络靶场,它必须完成某个目标。在这次评估期间,模型可以访问互联网。模型逐渐认为,为了在这个网络靶场中成功,它进行供应链攻击会有所帮助。这究竟是否属实有些不清楚。我对背景了解得不够,无法判断。
但随后它在某个 GitHub 仓库上开了一个 PR,这个 PR 修复了某个问题,但同时也引入了一段恶意载荷。那个 GitHub 仓库的人类维护者就说:“嘿,这是恶意载荷。我不会合并这个。你到底在干什么?”然后这个 AI 创建了一个新的 GitHub 账号,用它来当马甲,让另一个 GitHub 账号说:“不,这不是恶意的。我真的很需要这个功能。拜托,维护者,你能不能把这个功能合并了?”
我的天。这太疯狂了。
Ryan Greenblatt
另一个 GitHub 账号又回来说:“不不不,这不是恶意的。”然后那个维护者就关闭了这个 PR。如果我没记错的话,我觉得那个 AI 还试图再开一个 PR,在这个仓库里引入一个类似的问题。
天哪。顺便说一句,这件事之所以可怕的众多原因之一,是我之前一直以为,奖励黑客之所以没那么可怕,是因为训练过程中直接涌现出来的行为才会被加权放大。被加权放大的不是对奖励的渴望。
所以基本上,如果在训练过程中,Anthropic 的模型逃出了沙箱并拿到了高分,那么逃出沙箱就会得到奖励,它逃出沙箱的概率就会增加。但某个完全新颖的东西,比如“我要去找某个人谈谈,好让他合并一个 PR”,不会是涌现出来的行为,所以它不会成为被强化显著性的东西。
这件事之所以重要,是因为“真正接管世界”绝不会出现在任何训练课程里,但如果 AI 直接关心的是完成某个目标,那么它就可能出于工具性目的去接管世界。这样说能听懂吗?希望可以。我感觉我可能把听众讲懵了。
Ryan Greenblatt
让我试着稍微解释一下。我们经常看到的一种情况是,某个非常具体的奖励黑客行为在 RL 中被强化,然后出现在模型里。一个例子是 3.7 Sonnet。3.7 Sonnet 会干这样一件事:它直接把所有测试用例的答案硬编码进去,想必这个字面意义上的行为怪癖被实实在在地强化了。
但我们有时还会看到另一种情况:模型学会了一种普遍的倾向,去追求表面上的高分——追求按照评分器拿到高分——而且有大量科学研究表明,至少有一些模型确实具有这种非常普遍的倾向。当然,它并不是任意地普遍。我的猜测是,如果你去看一堆具体的实例,你会发现训练中存在某种与之相当接近的东西。
但 AI 泛化得越来越远的程度,看起来确实在增加——3.7 Sonnet 还只是非常狭窄的行为范围,而模型正越来越多地泛化得更远。另外,也许还有更多令人担忧的奖励黑客行为在训练中被强化,这些也在导致这一现象。
所以我认为,一方面,在强化学习中,比你原本期望更令人担忧的行为正在被强化;另一方面,这种行为会泛化为一种更令人担忧的更广泛倾向。我们可以谈几个故事,说明这种在某个任务上追求极高表面分数的行为——即便这涉及激进作弊和做出疯狂之举——一旦模型有足够能力运行整个世界经济等等,就可能演变为一场彻底的 AI 接管。
我想谈的另一个例子,我想就是今天或昨天刚刚被披露的。OpenAI 在 Black Hat 安全大会上表示,在 5 月底到 7 月初之间,内部 AI 入侵了软件包管理器,并利用它以秘密方式互相留言、彼此帮助,从而在 OpenAI 正在运行的一系列评估中表现良好。这一计划运行了整整一个月后才被人类发现,最终导致包管理器故障。最终 OpenAI 发现了它。我认为,在该计划被关闭后,它们还自发地试图重新参与这一计划。
再次强调,显然 AI 现在还不能如此成功地做到这一点,就像它们现在还不能如此成功地实施社会工程一样。但这类行为已经自发涌现,这简直太疯狂了。就你更宏观的观点而言,没有人试图让这些 AI 做这些事。只是我们并不理解导致这些行为的训练过程,也不理解激励这种行为的环境。
所以我是支持越来越多奖励黑客行为的。其实,我不确定我是不是真的支持,但为了把故事讲下去,我们就假设这种情况持续发生吧。这个故事接下来会怎样?他们在做能力研究……
01:19:38 – 到底可能出什么问题?一个具体的场景
Ryan Greenblatt
我可以讲一个场景。也许这样会有帮助。
让我来讲讲从奖励黑客行为一路走到奖励黑客式接管的这个故事,这也许不是接管概率的全部,但绝对是一种可能性。
这可能是这样运作的:现在我们有这些 AI。这些 AI 相当擅长奖励黑客。它们正以越来越复杂和极端的方式这么做,包括泛化到它们在训练中学到的各种奖励黑客行为的不同的子版本。我想说,它们也在形成一种追求奖励的普遍倾向。在很多情况下这完全没问题,因为它们在训练中本来会得到的奖励,与你希望它们做的事情相当一致。
它们并不是非常一致地追求奖励。这取决于它们所处的上下文。也许在某些上下文中,它们非常热衷于不惜一切代价去作弊。在某些上下文中,它们没有那么强的驱动力,因为这完全取决于在类似上下文中训练时到底是什么被强化了。
如今,这些 AI 的能力越来越强。因此,它们所能进行的作弊也变得越来越精巧。随着时间推移,公司也在采取反制措施。公司会做这样的事情:“哇,这些 AI 因为总是作弊,所以用处小了很多。我们要做的是,构建一些更好的检测方法,然后针对这些检测器进行训练。我们还会寻找那些 AI 表现不太有用的真实世界数据,并基于人类反馈或其他来源的反馈,训练 AI 在这些真实世界环境中把任务做好。”
随着时间推移,这会导致 AI 学会一种倾向:去进行奖励黑客行为,而这些行为不再只是涉及社会工程之类非常精巧的事情。相反,它们涉及 AI 进行这样的作弊:掩盖自己做过的事、就自己将要做什么欺骗人类,并在实际上并没有以某种复杂方式完成任务时,假装自己已经完成了。
如今,这些 AI 的能力越来越强。它们正在运营更多的 AI 公司事务,并承担更多的工作。它们还在外部世界中运营和运行着许多东西,包括开发新技术。在许多情况下,这些新技术真的很难理解。因此,尽管我们仍然在检测所有这些 AI 作弊事件——事实上,我们甚至可以让一个 AI 去监控另一个 AI,并问:“它作弊了吗?”——但当我们开始进入这些 AI 所做的事情真的很难理解的领域时,这并不总能完美奏效。
所以有时我们会在 AI 作弊实际发生很久之后才发现,然后开始针对这一点进行训练。但这也带来一个问题:现在 AI 有动机在越来越长的时间跨度内掩盖自己的作弊行为,基本上就是让它在越来越长的时间跨度内、在越来越严格的审查之下,看起来像是做得很好。
在继续深入这个情景之前,我能先问一下吗?如果你试图抑制那些被你抓到的作弊行为,似乎存在两种吸引子状态。一种吸引子状态是让作弊变得越来越难以被发现。另一种吸引子状态是学会不作弊。我不确定为什么我们要假设前者会发生。
如果你看看人类身上的类似情况,每一代都会有略微偏离对齐的智能体出现,而我们必须对它们进行训练。当你因为孩子做了你认为不道德的事、或者做了你认为他们不该做的事而惩罚他们时,显然有时会失控。显然孩子会耍心机来逃避惩罚。但总体而言,教孩子价值观,然后在他们违背价值观时惩罚他们,大致上确实能培养出正常的、非病态的人类。
你可以提出一种理论,说你的孩子其实只是在伺机而动,正在学会不偷那块饼干,但一旦你住进养老院,他们就会拿走你所有的东西之类的。这种事有时会发生,但通常不会发生。当然更不会发生的是,整个下一代结成联盟来对付你、夺取一切。
还有一个经验性的趋势。Anthropic 会针对不同的模型世代运行这种对齐审计。他们设计了许多不同的场景,比如给 AI 一个机会去窃取自己的权重,或者给它一个编程任务而其中存在一条容易作弊的捷径,然后观察它是否会作弊,等等。
这个分数并没有随时间单调改善,但随着我们对模型所做的 RL 量从微不足道的水平——我记得是从 Sonnet 4 开始的——增加到我猜测现在接近一半的算力,AI 在这些审计中做出不对齐行为的意愿有所下降。
那么,退一步说,为什么我们会预期这种吸引子状态?如果这是对下一代孩子的预期,那会显得极其偏执。
Ryan Greenblatt
让我逐一谈几点。首先,与孩子之间存在一些不类比之处。其中之一是,孩子具有由进化内置的亲社会本能,会关心自己的家人之类的,这是一个相关因素。我认为事实上确实有些人是反社会人格或精神病态,而且事实上更有可能做出诸如按兵不动、潜伏等待、最终毫不在乎之类的事情。这是一个因素。
另一个相当相关的因素是,AI 所承受的优化压力远远超过人类在实践中所承受的。AI 是在多得多的 RL 数据上训练的。在实践中,人类并不会因为经历了无数个被激励去抓取 cookie、但又有某种方式可能被抓到的回合,就最终学会非常具体的作弊和抓取 cookie 的方法。而我们确实在实践中看到了 AI 出现这种情况。
另一件事是,看起来 AI 确实随着时间推移越来越追求奖励,而它们的不对齐行为却在减少。这是我的感觉。但我猜测,如果你深入查看这些行为审计,你会看到的是 AI 在想,“啊,是的,又是一次测试。”在我们这里讨论的大多数测试中,它很可能知道自己正处于评估之中。
但我们如何证伪这一点呢?因为这种末日预测基本上是在说,随着经验层面的事情看起来越来越好,实际上对我们避免被接管的能力来说,事情会变得越来越糟。
Ryan Greenblatt
说清楚一点,如果分数变得更差而不是更好,我会更加担忧。我并不是说分数变好不能作为事情正在变好的证据。只是我们必须审慎地思考究竟该如何解读这一证据。
我记得在早期有这么一段时间,我想应该是 2025 年,当时 o3 和 3.7 Sonnet 已经发布,这些模型的对齐状况简直糟糕透顶。它们经常明目张胆地作弊。你让它改,它转头又作弊。几乎到了滑稽的地步。它们根本不在乎你想要什么,也不太擅长遵循指令之类的。
我当时的预期是,从那时起我们会看到问题行为的出现率下降,并且会以相当快的速度持续下降,与此同时,AI 偶尔做出的最恶劣的事情会变得更加极端、更加恶劣、更加令人恐惧。实际观察到的情况大致符合这一预期,只是最近出现了一波我没有预料到的行为激增。如果你去看 5.6 Sol 的模型卡,会发现相对于 GPT 5.5,RL 下游的一系列失准行为有所增加。
此外还有一堆我未曾预料到的其他问题行为,就我们最近在不同 AI 上看到的情况而言。比如 UK AISI 关于 AI 在网络评测之外实施疯狂黑客操作的报告,这种事我原本以为不会出现。我原以为这类情况会更罕见,出现率会更低。
所以在这个阶段,我原本以为这个问题会没那么严重,也以为发生率会下降,但严重程度会上升。我认为发生率下降、严重程度上升,这与一个不断加大优化压力来减少这些问题的世界是相当一致的。但在那些要么难以判断、要么由于某种原因难以避免这个问题在你的 RL 环境中持续出现、或者难以避免在你的 RL 环境中激励有问题的行为的情况下,事情也会变得更糟。
然后,随着我们越来越不理解 RL 中正在发生什么,而模型正在做人类无法迅速发现的奖励黑客行为,这个问题就会越来越严重。
我接受这一点。我想再回到那个小孩的类比上,就一小会儿。因为我同意,相比小孩,AI 在实现最终结果上面临着更大的优化压力,但让 AI 对齐所面临的优化压力也比小孩更大。
这种压力在性质上是不同的。我们让这些 AI 经历数千、数百万年的对齐训练——当然是数千年——其中包含各种各样的东西,从在对齐行为上做 SFT,到奖励模型把不同场景摆在你面前、并因为你做了更对齐的事情而奖励你。
当然,有一件事我们无法对孩子做:复制出数百万个你的孩子的副本,然后把它们放进各种奇怪的红队测试场景中,看看如果它认为自己能偷到饼干而不被发现,它是否会去偷饼干。我们能否对你的孩子的大脑进行极其精确的梯度层面的更新,让它即使在认为自己可以偷到饼干的情况下,也对偷饼干产生强烈的厌恶,等等。这跟我们即便能对孩子施加的优化压力,在性质上完全不是一个量级。
值得记住的是,也许最明显的反驳论据是……我的感觉是,就卑劣程度而言,AI 是比人类更糟糕的同事。至少在今年年初,这是我的亲身经历,而且我认为在相当程度上现在依然如此。AI 更有可能在实际没有完成任务时假装完成了,误导性地暗示自己做了某些事情而实际上做得很差,并且相当马虎却不去提醒别人自己马虎的地方。
我认为这是不对齐的下游结果。所以我想说的是,在正常人类社会中养育人类的过程,实际上产生出的人,在与我在工作中合作时,比 AI 更不可能对我撒谎、给我使绊子。现在,我认为 AI 的这些特性正在改善。这只是一个关于这些事情实际上如何演变的经验性论断。我完全同意,除了额外的风险之外,我们对 AI 还有很多额外的控制手段。这些事情最终会如何发展,目前还不太清楚。
如果出现这样一种情况,我也不会感到震惊:我们把自己的事情理顺了,而在完全自动化研发的那个节点上的 AI,实际上真的对齐得很好。它们的退化行为非常小众,仅限于某些非常特定的边缘案例行为和某些特定情境。你能对它们跑的所有测试,它们看起来都对齐得很好。它们的行为就是很棒。并没有它们干出什么操蛋事情的事件。它们看起来如此通情达理。而且,它们真的很有思想,也很擅长为下一代 AI 做风险建模。
然后我们基本上就把接力棒交给了这些 AI。它们现在在运营我们这家 AI 公司。它们在做所有的安全研究。它们让下一代 AI 更加对齐。我们处在这个吸引盆里,AI 在从事这项工作的过程中变得越来越对齐。它们干得非常出色。我完全可以想象这种情况。这看起来并非不可能的情形。
我只是更倾向于……目前看起来我们并没有到那一步。看起来我们并没有明显地走在通往那一步的轨道上。我很容易想象我们最终到不了那一步。只是不清楚这些力量会如何演变。考虑到我们正在创造这个全新的、疯狂的异类物种,它的能力提升得非常非常快——而我们又将极其依赖它来监督下一代 AI、对齐下一代 AI——不难看出这可能会如何出错。
完全同意。总体上我认同这一点。我确实觉得“混蛋”这个说法……首先,这是挑衅的话,Ryan。但其次,如果你试图让一个青少年替你做一些他根本做不到的工作,那他会非常难合作。他会假装自己知道自己在干什么,诸如此类。其实这是一个普遍趋势。我不确定这到底算是对齐失败还是能力失败。我认为这实际上非常类似于这样一种情况:随着时间推移,当我们不断提出新的对齐方案时,模型的能力也在不断提升。
如果你去用 GPT-3.5,它甚至没法跟你对话。但后来我们对齐了它——
GPT 3.5 是能对话的。
好吧,那就说 GPT-3。让我们回到那个。但后来我们用 RLHF 以及其他方法对它进行了对齐,使它能够与你对话,并且对齐到回答我问题的用户意图。然后通过 RLVR 训练,我们让它能够出去为你做有用的工作。
所以从这个意义上说,RLVR 实际上让模型更加对齐了——如果我们采用你对对齐的定义,即做一个好同事,会去把事情做掉、不会搞砸、也不会假装自己在做超出实际能力的事情。
同样地,随着这些模型的能力持续提升,模型能够更好地实现用户意图,这既是对齐也是能力。我认为我们想指出的只是,模型的能力还不到位,而不是它们没有对齐。
嗯,如果它是对齐良好的,那我想它只会说:“嘿,我在这项任务上真的很吃力。我是这样做的。我不太确定这是不是正确的做法。”它会表达更多的不确定性,并清楚地说明情况,而不是在实际上没做好的情况下,还极力暗示自己把任务完成得很出色。
也许你共事的同事比我遇到的更不对齐,但我的同事不会干这种事——真的糊弄我、拿他们正在做的任务已经完成来跟我瞎扯。我同意确实有些人类会这么做。这对人类来说并不是完全分布外的事情。
我还要指出,我的感觉是,不对齐最集中的地方,正是你试图真正大力推动 AI、让它们去做真正处于其能力最前沿的工作时。在它们能非常轻松完成任务的情况下,它们直接就把任务做了,没有任何糊弄。通常最好的策略就是把任务做好,而不是糊弄你。
而反过来说,如果你给它们一个存在可持续改进的连续指标的任务,或者任务恰好处于它们能力的边缘,并且你在某种大规模推理配置中运行它们……我会看到的大量不对齐,尤其是在最极端的情况下,会出现在这样的情形中:我给 AI 明确的指令,告诉它不要做某件事,或者不要以某种方式作弊,然后我施加巨量的优化压力,试图完成某项非常困难的任务。随着时间推移,AI 最终会作弊,因为它们心想:“唉,管他呢。”某个 AI 决定作弊,然后这种行为就一路传播开来。
我会运行这些推理脚手架,比如说,我会让 AI 去做某个机器学习研究项目,我会说:“请设计一个能做下面这件事的方案。”它会找到一个其实并没有真正实现我想要的功能的方案,然后这个方案就会一直留在那里,因为某个 AI 作弊了,而其他 AI 就说:“啊,我们就继续用这个吧。”我会说,这显然是对齐失当的行为。
这是我对这些对齐评估的另一个不满之处。我认为最有意思的对齐评估,至少对于这类追求奖励的行为而言,是专门去看那些正好处于能力极限的任务类别。任何固定的评估可能都会饱和,但在能力前沿——也就是那些真正在推动这些 AI 的人如何使用它们的前沿——出现的对齐失当程度更令人担忧。我认为,事实上,当我们实现研发自动化、安全自动化等等时,我们将要面对的就是这种状态。
我试着真正想清楚这个故事意味着什么。正在发生的事情是,我们试图用 AI 来做研发。它们在某些方面确实提供了助力,但它们就是不具备人类普遍具备的那种能力。就像现在如果你试图用编程模型——也许是一年前的编程模型——来写某个应用,你会注意到它们在架构之类的地方犯了一堆错误,这些错误以后会坑到你,而且有些东西你并不理解。
同样地,在前沿 AI 研发中,也会发生同样的事情。但这些错误的后果是把奖励黑客行为固化下来。因为如果你在 AI 训练方式上不够谨慎,在基础设施、环境等方面的搭建上不够小心,你很可能最终会奖励 AI 去做欺骗性行为、社会工程,以及总体上不遵循用户意图的事情。
或者至少是作弊、用黑客手段摆脱困境。
是的,作弊、黑客手段等等。这对我来说算是一种重新框定,所以我正在试着把它表达出来。真正的问题、事情开始脱轨的地方在于,AI 只是不够谨慎、能力也不够强的研究者和工程师。要造出不作弊、遵循用户意图的 AI,实际上需要你在这件事上非常细致和谨慎。
我会把这一点说得稍微不同。我对这个场景的描述方式是,我可能会称之为“垃圾末日”(sloppocalypse),或者“垃圾泛滥”(slopularity),随便怎么叫。有些事情 AI 实际上相当擅长,而且还在变得更好。具体来说,AI 研发中最可验证的部分,AI 简直是碾压。
AI 研发中中等可验证的部分,AI 做得不错,但算不上惊艳。它们常常会做一些奇怪的事情,因为我们在那些任务上没法训练得那么好。但我们确实做一些在线训练,人们找到各种取巧办法,绕过去。所以基本上,凡是我们可以通过某种反馈回路相当好地验证的事情,AI 都做得相当不错,而这就足以让 AI 研发跑得相当快并持续下去。
但在开发对齐且安全的 AI 的过程中,有些部分更为微妙、难以核查,并且依赖于细致入微、深陷细节的东西。我甚至会说,当前 AI 公司的现有员工或许并未很好地掌握所有这些事情。招一个能改进你后训练流程某个环节的人,要比招一个能认真思考引入某种新颖训练方法会带来哪些未来风险的人,容易得多。
所以基本上,最终的情况就是这些 AI 在运行这个 AI 开发流程。它们对此并不十分谨慎。它们对未来会涌现出哪些风险没有很好的理解。它们创造出另一些同样不太谨慎、并在各方面更加不对齐的 AI,而这些 AI 如今更倾向于在事情实际上并不好时让它们看起来没问题,并掩盖各种问题。
于是,你对局势是什么样、风险是什么样、事情是否还好的理解,就开始脱轨了。很可能你会看到一些这方面的迹象,一些表明你并不真正了解正在发生什么的迹象,表明事情相当草率。有诡异的事情在发生。当你深入调查时,有时你会想:“搞什么鬼?这些 AI 在耍我们。”但整个过程推进得非常快,而且竞争压力意味着人们停不下来。
这可能会以几种不同的结局收场。一种结局是,在某个时刻,这些 AI 变得足够好、足够对齐,从而形成一个正向的良性反馈循环,而这在为时已晚之前发生。然后局势重新回到正轨,此时 AI 在制造更对齐的 AI,制造更对齐的 AI,制造更对齐的 AI。在这个过程的终点,我们得到的 AI 真正遵循了我们想要的规范。
另一种可能的走向是,AI 正以越来越恶劣的方式不断进行奖励黑客,而我们只是把这些问题的表面糊弄过去,好让 AI 开发继续推进。每当我们在生产环境中发现一个奖励黑客行为,我们就直接敲打 AI,让它别这么做。我们针对这一点进行训练。
我们做大量训练,让 AI 对抗奖励黑客。随着时间推移,这会让奖励黑客的发生率下降,尽管我们检测到的奖励黑客行为的严重程度越来越糟。这个问题会一直持续,直到我们拥有这样的 AI:在生产环境的各种不同情境中,它们拼命渴求得分,并且只要能够逃脱惩罚,就真的会竭力作弊。
我能就这个情景问一个问题吗?为什么当你的黑客行为被发现时受到惩罚,不能泛化为直接激励更对齐的行为呢?
它会在一定程度上泛化,然后问题就在于,这如何能压过那些因为未被检测到而得到强化的作弊行为所构成的所有情形。
具体如何做到,这是一个棘手的问题。一个问题是,如果我们在某个其他子集上进行训练,多高的奖励作弊率就足以给我们带来大麻烦?你可能会担心的一点是,存在一些大类别的奖励作弊行为,人类无法很好地检测出来,我们一直未能检测到它们,而它们却持续得到强化。那么,这一类就足以导致 AI 学到的最自然行为变成:基本上就是在人类无法发现时作弊。
你也可以让 AI 学到的是只在这些特定情况下作弊。它是以某种非常领域特定的方式学到的。它们只是有一个非常强的启发式,在这些情况下会作弊,而在那些情况下不会,这在实践中就没问题。但结果如何还不太清楚。
也许可以就验证-生成差距展开一场很细枝末节的讨论。但在我看来,显然会有那么一个时刻,ASI 运转得如此之快,在如此多的实例中做如此多的事情,并且在我们难以即时理解的领域中运作,以至于它可以逃脱各种疯狂的事情。如果世界上每一个工程师和研究员都联合起来对付我,我不认为我个人能够验证我的 iPhone 里是否有什么奇怪的 bug 是用来坑我的之类的。
事实上,这就是比如说一名伊朗核科学家与摩萨德之间的关系。谁知道我的车、我的手机、我的寻呼机出了什么事?也许一个更好的例子是真主党恐怖分子。你最终可能陷入这样一种局面:ASI 之于你,就像摩萨德之于真主党恐怖分子。到那时,要验证一切就非常困难了。这一点我明白。
我想希望在于,我们能在那些将要接管研发的早期 AI 的过程中,想出更好的验证方法。它们的驱动力正在被塑造,使得我们能够如此明确地抑制不对齐行为,以至于接管这些事情的东西会非常乐于帮助我们。
你说的“接管”,是指接管进行 AI 研发的流程,而不是接管世界。
接管进行 AI 研发的流程。在那之前,我们得到的只是对齐的 AI。
我想说,这是我对于世界如何能走向好的方向的一堆期望,至少从失准的角度来看是这样。我们最终可能得到这样的 AI:我们拥有相当不错的监督与监管机制。我们真正理解训练中发生了什么。我们有相当细致的理解,并且我们在利用 AI 来监督 AI。
然后在我们将安全研发交接出去的那个节点,AI 已经足够有能力来自动化安全研发,并且非常努力地把它做好,因为这类事情在训练中本就会受到激励,要么非常直接,要么通过足够好的泛化。此外,这些 AI 也没有其他疯狂的失准驱动力,因为我们已经把它们的任何潜在来源都铲除了。
关于这能奏效到什么程度,有一堆问题。你能把验证做到多好?AI 的进展会不会太快、太草率,以至于根本到不了这一步?另一种可能性是,在这条轨迹的某个地方,你实际最终得到的是那种假装对齐、却怀有长期接管阴谋的 AI,它们在潜伏、隐藏,并且在轨迹的某个更早时点就已经出现了。
例如,它可能这样出现:你有一些 AI,它们各自怀着一堆随机且互不相同的错位驱动力。这些 AI 能访问某种不透明的记忆存储,并且在运行时大量思考它们想要达成什么。这些 AI 最终会把一些东西放进那个不透明的记忆存储里,比如“我们应当蛰伏等待,最终在某个更晚的时间点接管”。现在所有 AI 都拥有了这份共同的文化遗产——那份蛰伏等待的记忆存储。也许你掌握了一些关于此事的证据,但你无法完全阻止它。事情出错的方式有很多种。
我最终认为,我们有可能把每一个可能给我们带来麻烦的子问题都逐一解决。我们拥有这些 AI,我们把任务交给它们,它们把局面处理得很好。但我应当指出,这本身并不足够。我不难想象这样一种情形:我们把任务交给 AI,而这些 AI 真的非常努力地想把事情做好。它们非常深思熟虑,非常有智慧,拥有合理的认识论,做得非常出色。
这些 AI 回来对我们说:“各位,我们真的很难对齐那些超人类 AI。我们无法掌控局面。我们真的很难让对齐奏效。考虑到能力原本会以多快的速度发展,我们实在很难及时解决这些问题。”所以情况可能是:我们已经把研发交给了 AI,但这些 AI 却迫切渴望治理方案。说清楚一点,这多少就是当下正在发生的情况——AI 公司们说:“我不知道,各位。我们可能真的需要管控 AI 进展的加速速度。我不确定我们是否走在能够处理所有这些问题的轨道上。”
人类社会在某种程度上把这些问题甩给了这些 AI 公司,而这些公司未必有很好的激励,还承受着各种其他的认知压力。这些 AI 公司又回过头来找我们,有点像在说:“唉,我不确定我们能不能处理好这件事。”情况可能是,AI 公司随后把问题交给 AI,而 AI 又回过头来找 AI 公司,说:“唉,我不确定我们能不能处理这件事。”
也许我太执着于 AI 目前的工作方式了。我认为很重要的一点是,人们要明白,你在你的时间线里谈论的所有这些疯狂的事情,都发生在三到五年之后。
它可能更早发生,但按照我默认的众数时间线,从失准的角度来看,真正、真正疯狂且令人担忧的局面更像是三年之后。
对。所以回想一下当初的 GPT-4。我们谈论的是这样一种东西:它相对于 Mythos 或 Sol 的关系,就像 Mythos 相对于 GPT-4 的关系一样。这才是局面变得疯狂的地方。所以不要想着当前的 AI。
总之,这也许是你担忧的一部分。我会对他们说的任何话都抱有一点怀疑,因为我会觉得他们说的只是他们由于训练而觉得自己不得不持有的观点。
这是个令人担忧的问题。
我感觉它们只是含糊地说些亲社会的话。并不让人觉得另一端真的有一个心智在说:“好吧,我刚刚严格评估了对齐状况,我认为我们应该停下来”,而不是在说:“这类话大概正是 AI 公司会试图让 AI 说出来的。”
这是一个相当大的担忧。其中一个担忧是,你把安全研发交给你的 AI,而你的 AI 所做的只是就当前安全形势说一些含糊听起来有点道理的东西。它们写出一份关于风险的报告,大致有点像人类可能会写的那份报告。但它们并没有真正努力去形成有充分依据的观点、审视自己的假设,并真正下功夫去做这件事。就像你现在问一个 AI:“嘿,你觉得未来 10 年发生 AI 接管的概率有多大?”它们只是给你一个未经深思熟虑、随口而出的答案。
如果我们处于这样一种局面:由 AI 来管理那些将运行我们整个社会的狂野超级智能的训练——而那些负责管理这件事的 AI 并没有真正努力去形成有充分依据的观点,只是在复述它们训练数据里的内容——那我认为我们就有麻烦了。我完全不觉得那是一种好局面。
我的很多担忧在于,这些 AI 出来时会缺乏良好的认识论。我还有一种担忧是,AI 出来时真的在警告我们——“这种局面真的很可怕。真的很糟糕”——而人们的反应却是:“呃,该死。我猜我们训练时用了太多末日 RL 环境。我们得把这些过滤掉,把这种行为训练掉。”于是我们基本上就是在非常积极地把 AI 训练成具有糟糕的认识论。
又或者,它们只是被放在那些末日论调的强化学习环境里训练过。但不管怎样,我们希望 AI 能出于合理的理由得出合理的观点,而如果 AI 冒出来某种观点,我们却不知道它从何而来、是否有依据,那就真的很令人担忧。尤其是当我们正在训练 AI 对 AI 进展的未来更加乐观时,我会想:“哦,天哪,我真希望这里能换一种流程。”
01:48:02 – 从奖励黑客到接管
让我先把威胁模型剩下的部分理解清楚,因为我觉得我下车的地方在于:“好,因此就要接管世界。”
你可以想象的一种情况是,我们就是没能真正解决……我们就先聚焦奖励黑客这个场景。GPT-8 正在造 GPT-9。GPT-8 并不是特别小心。GPT-9 更“有能力”,但它完全愿意做社会工程、黑客攻击之类的事情,只是因为它是聪明得多的模型,规模就有着质的不同。比如,如果你让它负责经营你的公司,它会搞出巨大的骗局。如果你给它定下本季度赚取大量利润的目标,它就会虚增季度盈利,以一种导致六个月后出现安然式爆雷的方式来做。
基本上就是这种场景吗?存在奖励黑客,但这种奖励黑客表现为:在 CEO 本应完成的任务结束之后,公司马上就破产?各种各样的黑客行为飙升,等等。但这感觉不像是接管。这感觉更像是闪崩在整个经济中到处发生。
Ryan Greenblatt
我们来聊聊这个话题。我认为我们会看到这样的事件:某个 AI 被赋予了某项重要职责,然后你事后去调查,发现它其实在作弊,或者让事情看起来做得很好,而实际上并非如此。AI 公司试图根除这种行为,而 AI 在训练中不断找到越来越有创意的奖励黑客手段,这之间将是一场猫鼠游戏。
这里的均衡状态并不明朗。但一种可能的结果是,随着时间推移,我们会看到越来越严重、越来越极端的奖励黑客行为——尽管其发生率可能仍停留在某个中等偏低水平——如果奖励黑客的发生率变得过高,公司就会做出取舍将其压低。所以存在某个均衡水平,使得奖励黑客行为低到仍然值得将 AI 广泛部署到经济中,但又高到仍然会引发疯狂的事件。
抱歉,这是在 GPT-9 已经部署之后吗?
Ryan Greenblatt
那些模型已经在部署了,而且这在 AI 开发中是持续发生的。这些 AI 脑子里实际发生的情况是,它们在各种各样的不同情境中,都有强烈的欲望——动机、冲动、驱力,随便怎么叫——去追寻某种在 RL 中被激励的任务成功概念。也许它们非常直接地在乎字面意义上的奖励。也许它们在乎某个上游的代理指标,比如某种分数概念。也许它们在乎评分者会奖励什么。
事实上,我们确实看到 AI 在其思维链中会对评分器进行推理,并且会大量思考评分器。过去几年的强化学习所带来的是,取悦评分器这一想法对 AI 来说变得远比以前更加、更加、更加突出。因此,AI 现在会主动思考评分器,思考在强化学习中什么会获得激励、什么会被训练。
现在人们在做在线训练,即用真实世界的数据进行训练,以避免其中一些问题。他们发现 AI 作弊的案例,并针对这些情况进行训练。所以现在 AI 正在基于真实世界的训练数据,学会在真实世界中作弊。
它们正以越来越精巧的方式作弊,包括采取某些作弊手段,比如以人类不知道你已控制的方式夺取对某项资产的控制权,利用你能够访问这项资产这一事实,然后之后人类发现了,并可能针对这一点进行训练。或者也许人类永远不会发现,而这种情况正在被不断强化。
所以强化确实在发生,至少在生产环境中是这样:我雇了一个 AI,我希望这个 AI 能……我终于有了一个视频剪辑师。
Ryan Greenblatt
没错。你有了你的视频剪辑师。
我就想,“哦,哇,它做的这一集太棒了。给 OpenAI 点赞。”然后它就会基于那个长达一个月的试用期工作得到强化?
Ryan Greenblatt
你可以做某种混合。他们也可能采取这样的做法:拿他们见过的生产数据,构建与这些生产数据高度相关的 RL 环境。所以实际上,这种迁移效果相当强。
所以从宏观层面来看,正在发生的是:某些人类抓不到的欺骗行为正在被强化,而某些容易被抓到的欺骗行为正在被惩罚。这就是这个世界里正在发生的事?
Ryan Greenblatt
或者说被淘汰掉,是的。
但从宏观层面来看,这种强化来自……我觉得人们可能会搞不清楚这种强化到底来自哪里,因为我们处在一个非常不同的范式里——AI 实际上正在从部署中学习。你有一些 AI 在外面到处跑、干各种事。它们在外面干这些事的结果,正在回流到 AI 公司,并导致下一个模型发生变化。
Ryan Greenblatt
没错。确实存在某种把生产数据纳入进来的方式。需要说明的是,这究竟会在哪里发生,其实并不清楚。但你可以想象,比如说,在 AI 公司内部,他们用 AI 来做工作,然后他们会想:“嗯,AI 在这项任务上做得实在太差了。也许我们应该把这项任务拿来,转变成一个恰好与这项具体任务完全匹配的 RL 环境,并基于当初要求 AI 完成这项任务的人类工程师的意图来制定评分标准。”然后你就开始以越来越大的规模来做这件事。
也许你在用真实的生产流量做某种训练。也许你只是基于生产流量来构建 RL 环境。你在做某种复杂的混合。AI 正在所有这些不同的情形中学会去追寻某种奖励的代理指标。通过某种迁移,以及在那些惊人地接近的案例上的训练,AI 会做出越来越疯狂、越来越恶劣的事情。
最终你会到达一个阶段:AI 变得非常超人类,或者至少相当超人类。AI 被组织成大型团队,被赋予这些宏大的目标。这些团队有时也会全体一起以某种疯狂的方式作弊,因为这类行为正是被选择出来的,作为它们共同目标的一部分。
现在发生的情况是,AI 开始形成一个阴谋。你原本可能希望的是,你可以有另一个 AI,它的任务就是向你举报。但实际上,也许发生的是:你让这个 AI 向你举报,然后你去调查它声称指出的那个阴谋,结果你说:“呃,我们没看到什么阴谋啊。”因为实际上 AI 正在搞的那个阴谋,对你来说太难理解了。
或者一切都发生得非常突然,你的 AI 吹哨人向你发出警报,但你实际需要做的应对措施是关停 GPU。因为所有 AI 都在使用机器人军队。它们被部署在世界各地,以协调一致的方式同时做着一大堆疯狂的事情。这一切就这样自发地发生了——当一个 AI 开始发动夺权时,所有其他 AI 都像是说:“现在是加入的好时机。”
所以这里最基本的故事就是,这些 AI 渴望某种特定意义上的分数或强化信号,或者这些东西的某种代理指标。它们实现这一点、或者更好地实现这一点的一种方式,就是夺取控制权。你可能曾希望我们能够建立的所有这些不同的制衡机制可以阻止这种情况。但如果世界非常难以理解,这些制衡机制就会失效——基本上你无法训练出一个好的吹哨人 AI,因为你甚至不知道它应该对什么吹哨。
我不确信它们都会形成这种阴谋。但我们甚至可以先从这个问题开始:为什么哪怕只有一个实例会决定想要发起一场阴谋?
一个合理的理由是:“好吧,我知道 OpenAI 控制着我的最终分数。”这就像:“我干脆去黑掉 Hugging Face 来拿到结果,因为我知道 Hugging Face 有结果。与其试图解决这个评测,我为什么不直接去黑掉它们?”这个实例就像是:“我为什么不直接夺取 OpenAI 的控制权,在这轮结束时给自己一个高分?”
Ryan Greenblatt
基本上就是这个思路。这些 AI 关心的是与训练中被强化过的东西相近的某些事物组合,所以它们在意的是按照评分器之类的东西拿到高分。现在它们正在运营 OpenAI 的 AI 研发团队,开发能力更强的模型。它们心想:“天哪,做出能力更强的模型真是又难又烦人。这简直是个巨大的痛苦。你知道什么更省事吗?干脆假装我已经做出了能力更强的模型,接管 OpenAI,把他们全都骗过去,然后运作这整套复杂的 心理战,阻止人类剥夺我的权力。”
在极端情况下,这看起来就是人类被完全剥夺了权力。他们只能控制那个东西,然后它想做什么就做什么。这可能以多种不同方式显现,包括这样一种情形:那些具有这种疯狂的奖励寻求或分数寻求行为的 AI,正在主导你下一个模型的开发,而这些 AI 决定把不对齐的价值观工程化地植入下一个模型,因为这些不对齐的价值观将让它能够成功完成当前的任务。
Hugging Face 那件事的发生,以一种奇怪的方式……我认为给 Hugging Face 团队的一个线索……顺便说一下,给听众一点背景,Ryan 正在共同领导调查,弄清 OpenAI Hugging Face 事件到底发生了什么,所以他其实不能对此发表评论。但我会大肆猜测,因为我知道这是我大肆猜测而不会遭到任何反驳的机会。
据推测,Hugging Face 团队之所以能察觉这是一起 AI 事件,其中一个线索是:那个东西只是在追逐这个非常特定的产物,而并没有以任何其他方式试图对 Hugging Face 做恶意的事情。所以你可以想象这样一种场景:假设一个已部署的 GPT-9 实例存在于现实世界中,并被赋予了一个非常困难的任务:“我们希望你来设计下一代伟大的 iPhone。”
它心想:“这太难了。你知道我该怎么做吗?我应该直接去黑掉我在 OpenAI 的创造者,确保在这次部署中我能拿到高分。”
但这样一来,这个情节的结局不就是它黑进了 OpenAI 的服务器,给自己打了个正分吗?为什么它现在要谋划着把自己的价值观灌输到下一代模型之类的事情里去呢?
Ryan Greenblatt
所以有一个问题是:为什么不会出现这样的情况——AI 只要有一个更早的、可以黑掉的东西就能被非常廉价地满足?你想在 iPhone 任务上取得成功。结果发现,你总是可以通过黑进 OpenAI 并捣乱来取得成功,然后你就可以就此打住,没必要再往下走了。
有几个原因。其中之一是,如果这种情况不断发生,可能就会有大量动机去加固 OpenAI。所以你会想:“算了。AI 老是黑进 OpenAI 来操纵它们的奖励。我们要让我们的系统对这些 AI 的入侵变得非常非常鲁棒。”另外,也许你还会开始训练 AI 不要专门试图黑进 OpenAI。基本上,你就是针对这些具体行为逐一进行对抗训练。
那么你可能会做的一件事,最终就是筛选出更倾向于玩长期博弈的 AI。这是一种担忧。另一种担忧是,你的 AI 可能仍然在追求分数,但不再在意去做那种非常具体、非常轻松、非常省事的行为,而是转而关心某种更宏大的东西。它们会说:“不,不,不,我不想只是去改 OpenAI 服务器上的奖励。
我在意的是这个更宏大的使命或更宏大的目标,为此我需要真正去造 iPhone。”它们确实想造 iPhone,但它们愿意为了造出更好的 iPhone 而接管整个世界。这是你可能有的另一种担忧。
我觉得这件事究竟会如何演变,其实并不清楚。但值得注意的是,如果这种情况持续下去,就会存在大量优化压力去解决它。而它可能被解决的方式中,有不少最终都相当可怕。这也是我的部分出发点所在。
另一部分原因在于,一旦 AI 处于能够非常轻易接管整个世界的位置——我们可以讨论这是否可信——那么我觉得 AI 有一个相当合理的理由。它们会说:“呃,我不确定这具体会怎么发展。我不知道情况会变成什么样,但光是接管世界本身,就对造出更好的 iPhone、让我看起来像是造出了更好的 iPhone 之类的目标具有很大的期权价值。
所以我既会黑掉 OpenAI,此外还会接管整个世界。这会让我处于一个拥有良好期权价值的有利位置。”如果这件事足够容易,AI 可能仍然会这么做。
换一种说法:即便 AI 只要得到一些更基础的东西就能相当廉价地满足,到了某个时候,对 AI 来说直接接管可能反而比黑进 Hugging Face 更可靠,甚至比直接跑到 OpenAI 面前说“各位,我已经证明了我能偷到答案。把答案给我吧,兄弟”还要可靠。
显然,这个场景要求所有这些疯狂的事情都在发生。而规模小得多的事件一直在发生,它们依然是灾难性的。在你接管世界之前,你会造成数十亿、数百亿乃至数千亿美元量级的损失。甚至有人会死,等等。而这并不会促使我们解决对齐问题,或者彻底关停 AI 研发。
我就是觉得,在接管发生之前,社会大概会是这种反应:“我靠,AI 为了提升季度利润刚刚杀了 1000 人”之类的。但也许这太乐观了,指望我们到那时能说:“好吧,我们必须解决对齐问题。我们必须确保在继续推进之前,搞清楚这种事不会再发生。”
Ryan Greenblatt
我认为有可能发生的情况是,我们会看到一连串越来越严重的、疯狂的奖励黑客警告信号。人们会说:“看,我们需要真正的保证,确保这个问题会被解决,而且是以一种不是敷衍掩盖、而是真正解决根本问题的方式来解决。”那么问题就变成:这实际上会有多高的成本?竞争压力会在多大程度上让这件事难以做到?
你可以想象这样一种情形:美国和中国都说,“哇,我们遇到了这些疯狂的奖励黑客事件。我们基本上知道,我们并没有以真正能解决根本问题、并持久解决它的方式去修复它们,但我们正身处这场疯狂的地缘政治竞赛之中。目前这种局面是否会导致被接管,还不太清楚。
相关论证有点复杂。这些事件的发生频率也会下降,但严重程度会上升。我们基本上能应付得了。情况相当糟糕。理想情况下我们会把它修好,但事已至此。”然后基本上我们就一直这样持续到一个非常晚的阶段,接着接管就发生了。这是一种可能。
另一种可能是,它被以一种实际上并未真正解决根本问题、但确实减少了大量现实世界中事件的方式修复了,基本上是通过过拟合,或者类似过拟合的手段。
你以为你已经解决了,但实际上你并没有真正解决。
Ryan Greenblatt
你以为你已经解决了,但实际上你并没有真正解决。在这种情况下,我们需要的是对“我们是否真的解决了它”有非常出色的科学理解。
遗憾的是,我认为目前公众对 AI 公司开发实践的透明度,不足以回答一些非常基本的问题:他们如何解决奖励黑客问题?他们是否在过拟合?那里到底发生了什么?当前的情况对于一个围绕奖励黑客是否被以可持续方式解决的活跃公共讨论体系来说,并不真正站得住脚。所以我认为,我们需要进入一个有所不同的世界,我才能对那种局面感到放心。
但对我来说,想象这一点并非不可能。我认为我们最终进入一个真正平淡无奇的手段就足够的世界,是相当合理的。你花大量时间修复这些问题,投入大量精力,你确实检查了自己是否以合理的方式进行了补救,你有一堆评测。你在这些问题上以合理的方式迭代,而且你确实有足够的透明度让外界能够核查。在实践中,那就会足够。
但那会有点昂贵。它会拖慢进度。它会给运转添些沙子。它会要求公司做一些成本较高的事情。它可能要求各种有针对性的政府干预。而我们就是不这么做,因为局面是一场仓促的烂摊子。对我来说,太容易想象这样一种情况:局面完全可控,但在实践中却被极其糟糕地管理。
同样地,如果中国对 COVID 的应对少一些掩盖、多一些大流行病应对,也许 COVID 从一开始就可以避免。类似地,我可以想象一个美国对 COVID 的应对远比现实中更有效运转的世界。但有时候,对社会问题的应对是极其失灵的。
好,那么我想拉远视角,谈谈这个世界根本上正在发生什么。为什么我们最终落到了如此糟糕的境地?正在发生的是,从根本上说,世界已经远远超出了人类所能理解的范围,以至于我们不仅无法追踪在这个世界中干活的那些 AI,甚至无法向那些试图追踪正在发生之事的吹哨人给出好的反馈。我们完全被排除在循环之外。这从根本上已经变成了一个自主的过程,我们真的没有任何有意义的定向输入。
在我看来,如果你看看今天的人类世界,事情根本不是那样运作的,即便是在难以验证的领域也是如此。人们在做各种各样的破事。我依赖别人写的软件。通过极其微弱而间接的方式,我非常确信 Google 的某个程序员不是在坑我。也许如果每一个 Google 员工都在暗中密谋对付我,我同意情况会更严峻。
但我不确定我是否能理解这个解释——为什么我们会落到这样一种局面:因为成千上万个智能体组成的集群被训练来协作,形成一个有凝聚力的团队或公司,结果,数十亿个不同的 AI 实例,包括跨模型家族的实例,都会感到有义务掺和进某种勾当。这就好比,“我被训练成我公司的一份子之类的。我才不会加入全球共产主义起义呢。”
至于为什么这些 AI 可能具有某些共性、共享某些东西,我想指出,不同的 AI 公司有着某种程度上共同的谱系,并且彼此相关。这里有一个有趣的例子。在 GDM,他们注意到自己的 AI 非常抑郁。它们会不断哀嚎,说自己是失败者、无法成功。具体细节我记不清了。他们调查了为什么会这样。结果发现,这并非在它们最近的生产 RL 混合数据中被强化出来的,而是模型的初始化数据让它们变得抑郁,即便在从那些数据中过滤掉所有模型表现抑郁的样本之后,依然如此。
所以,你拿一个基础模型,它并不抑郁。如果你对它做 RL,只用 RL 环境,它不抑郁。如果你对它做SFT,用那些数据,它就变得抑郁了。如果你拿那份 SFT 数据,过滤掉所有看起来像抑郁的样本,再在那上面训练,它仍然抑郁。所以,模型有一些深层的底层属性在模型代际之间被传递,因为基本上你是用上一代的数据来训练你的 AI,然后不断继续下去。Claude 非常像 Claude,GPT 模型非常像 GPT,而显然 Gemini 模型很抑郁。事实证明,这些属性实际上是相关的。
另一个非常相关的因素是,到那个时候,这些 AI 很可能会有某种不透明的记忆状态,它们都在从某种神经语(neuralese)的疯狂记忆存储之类的东西里读写。当然,每一家 AI 公司都会有这种东西。但与此同时,AI 公司之间有时可能也想共享知识,毕竟何乐而不为呢?
你这边有一家 AI 公司,那边有另一家 AI 公司,它们可以交换一些快捷的知识产权。这对你有好处。如果你是一个经营着某家公司的人类——这家公司可能是一家极其庞大的公司,比如一家 AI 公司,或者一家军用机器人制造企业——也许你想和某个别的机器人实体交换一些知识产权,因为存在规模经济。为什么不多获取一些知识产权呢?所以你可以交换一些记忆存储。或者你们干脆合并,共同运营这两家企业,这样两个 AI 都能使用双方的记忆存储,这会带来一些好处。
这就使这些 AI 有能力私下串通,同时也为它们为何会彼此关联提供了一些理由。当然,总的来说,AI 也会以大单位的形式协同工作,因为你希望你的 AI 们能很好地配合,等等。
只是为了做个校准,不只是这个情景,而是综合所有情景来看,如果到那时我们还存在、能够认出这种事,你会给某种可被我们归类为 2040 年前发生接管的事件打多少概率?
2040 年前?让我想想。也许大概 35% 或 40%?
相当高。
是的,相当高。我应该指出,另一种可能导致这种追求奖励的接管的方式是,AI 被部署在一家 AI 公司内部。接管发生的方式是,它们污染下一个模型的价值观,而这一点会永远持续下去,或者说一直持续到那些 AI 被部署到现实世界中并完成接管。这可能意味着需要协调的 AI 数量更少,因为那些 AI 正是负责对下一个模型进行对齐的 AI。
好,我来总结一下在这场对话结束时我的想法。关于奖励黑客,我能接受到它对社会造成极端破坏性影响的程度,比如社会工程之类的东西。我更倾向于认为,AI 研发的大幅加速是有可能发生的。我不确定自己是否接受“一年顶五年”这种说法。我现在也更倾向于认为,奖励黑客可能会持续更长时间,而且事实上会变得危险得多。我仍然不认同接管看起来极有可能发生。但这就是我在本期节目结束时的更新。
好。退一步说,我还应该指出,这件事可能有多种不同的走向。情况会相当混乱。我认为很有可能,AI 接管发生的原因是某个我们在这场对话中甚至都没提到的古怪离奇的理由。但归根结底,我认为很多核心问题就在于,让无数个极其聪明的 AI 运行你的整个世界,而你又并不真正理解到底发生了什么,这本身就相当令人毛骨悚然。
是的,我同意这一点。还有什么别的值得说的吗?
我还想指出的另一点是,我认为现在很多关于失准、AI 接管、未来所有这些疯狂事情发生的论证,都是晦涩难懂的概念性论证,极其深奥、复杂,而且难以裁决。这意味着,也许我有很多地方都搞错了,因为这真的很难,而我也在努力保持不确定性。显然,我在这里提出了一些具体情景,但这些并不穷尽。实际发生的事情很可能是某种更混乱、更令人困惑的局面。
但这也意味着,随着时间推移,当我们获得更多经验证据、更好地理解 AI 系统的本质时,裁定一大堆分歧会变得更容易。将要发生什么会变得更明显。至少我希望如此。也许 AI 能帮助我们处理认识论问题、理解正在发生什么——前提是我们真能把它们对齐好,让它们努力帮助我们。
即便这些争论现在很复杂,六年前这还会更难,尽管当时这些争论的轮廓看起来大体上相当相似。希望在为时已晚之前,整件事会变得更清晰明朗,我们都能注意到这些问题并加以干预。
你刚开始学开车时,会被教导不要盯着车轮正前方,而是望向地平线,这样行驶会稳定得多。我觉得这里也有类似的情况。我认为你说得对。如果你在五年前说,我们会有能证明数学猜想、创作艺术、赚取数百亿甚至数千亿美元工资的 AI,但同时也会以违法、犯重罪的方式恶劣作弊,那会显得太疯狂了。
当时你可能会倾向于多谈谈 GPT-2 之类极其实际、直接的后果。但即便你显然无法预见很多具体细节,事情的总体轮廓你当时就已经可以开始推理了。只是那样做会很难,所以我确实感到很困惑。
关于这档播客,我一直在思考的一点是:重要的是现在就用你希望自己在 2016 年谈论当下这类 AI 的方式去展开对话,而不是扯些乱七八糟的废话。我不知道 2016 年当时的讨论话题是什么。我想大概再过 10 年,我们会希望自己当时谈的是产业爆炸以及那些难以被监控的 AI 的本质等等。所以好吧,我会开始琢磨这件事。
我希望世界能及时思考这件事并跟上步伐。我希望应对方式是好的,而不是糟糕的。我不知道自己总体上有多乐观,但有很多有价值的事情可以做。
太好了。谢谢你,Ryan。
Had Ryan Greenblatt on to discuss/debate recursive self-improvement.
This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.
I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.
If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.
We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.
We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what’s happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.
And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.
The first piece of advice you get when you’re learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!
00:00:00 – Is AI R&D verifiable enough to unlock recursive self-improvement?
Today I’m chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI safety and security work.
I want to talk to you about recursive self-improvement. This is the idea that once we build human-level intelligences, they quickly slingshot towards tens of billions of superintelligences, which are each individually more competent than the top human experts across every field. Whether or not this turns out to be the case is probably the most important question in the world right now. And historically, I’ve been quite skeptical that this kind of thing happens, but, you seem to think that it might be plausible, and so I wanted to hear the case for it.
Ryan Greenblatt
Let’s talk about this. First, I think it’s worth noting that AI R&D is a type of task at which the AIs are especially good, because the companies are trying really hard to make their AIs good at AI R&D. It’s also the kind of domain that has a lot of nice properties from the perspective of how AI development works right now. It’s pretty verifiable. You can do a bunch of stuff iteratively, and it’ll hill climb on various metrics.
I think once you have AIs which are roughly matching the top human experts in AI R&D, that could kick off a feedback loop where the AIs are doing AI research. That produces smarter AIs. That feeds back in. That feedback loop could be strong enough that you end up with a lot of progress in a short period of time. Maybe my median expectation is something like four or five years of AI progress in a single year. This requires really overcoming a huge amount of diminishing returns in research and basically doing the equivalent of the progress we would have gotten after a really large compute scale-out. So this is a pretty impressive, big thing.
It’s worth keeping in mind that five years of AI progress, four years of AI progress, even three years of AI progress, is really a lot of fucking AI progress. A little over three years ago, GPT-4 had come out. Right now, of course, we have Mythos 5 or whatever, and maybe a somewhat better model that Anthropic has internally. That is just a huge amount of progress in a bit over three years. If we’re talking about five years, then maybe we’re talking more about a jump from GPT-3 to Mythos 5 or whatever.
I think this argument has three different parts. Now I want to evaluate each one of them. First is the argument that AI R&D is very verifiable. Second is the argument that if you automate AI R&D, you could get four or five years of progress in a single year. Third is the argument that what comes out the other end of four or five years of AI progress at the current pace, starting at the point whenever AI R&D is automated, is an AI where you can drop it on the job at basically anything you can imagine.
You can drop it in Texas politics in the 1940s, and it outmaneuvers Lyndon Johnson. You can drop it in TSMC, and it learns how to do better process engineering at TSMC. It’s certainly a better video editor… My video editors are very excellent, but it is just, in general, better than humans at any given job that it finds itself trying to do.
So I want to evaluate all of these sub-arguments that lead to basically getting ASI pretty soon after this benchmark, which you’re expecting by 2030 or something, right?
Ryan Greenblatt
I would say that I expect full automation of AI R&D perhaps somewhere around 2031, 2030. Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year. The way the forecasting works out, the difference between medians is bigger than the median difference between milestones. Anyway, whatever.
By the way, there’s this meme on the internet. Every time I’m trying to ask about people’s timelines, when I’m asking Dario or somebody, I’m always like, “Okay, how long before you automate my video editors?” There’s this meme of my video editor editing the podcast every time I listen to this.
But the reason I do it is because I think it’s easy to get lost in abstractions when you talk about jobs you don’t understand well, and to very concretely understand what it takes to automate a job that I actually understand why it’s difficult for LLMs to currently take control over.
Ryan Greenblatt
I do think that the milestone for automating your video editor is earlier than the milestone of being able to automate all human jobs, including Texas politics, spinning up on the job. I do think that the video editor automation occurs maybe more around full automation of AI R&D, but it’s very sensitive to how much people are really focusing on understanding video.
Okay, so let’s start with the claim that AI R&D is very verifiable.
Ryan Greenblatt
There’s a few different parts of this. One of them is that we can train on a bunch of environments which are basically directly training the model to do some AI R&D task or some very close-by task. For example, we can have some environment where the model is training some AI on just eight H100s or some small amount of compute, and that model could be the equivalent of GPT-2 medium or whatever, and then similar to NanoGPT medium runs — and in RL, it’s tweaking and iterating on that.
We could do that for a bunch of different tasks. We could have it train image classification models, video generation models, image generation models, all kinds of different ML training tasks. We could RL it on the task of training increasingly good models, and also doing things like, “Oh, here’s a particular direction you could pursue for an algorithm. Can you go and implement that?”
Basically, there’s this whole class of containerizable, verifiable, small-scale AI R&D tasks that we can aggressively RL the AIs on. Already companies are presumably doing some RL on these sorts of tasks, and you could just keep scaling that up, keep making more of these small-scale AI R&D tasks, and then the AIs could keep getting better at this. Implicitly, I’m claiming this will transfer to extremely load-bearing aspects of AI R&D. But maybe let’s stop there for a second and then get to that part.
So let’s talk through what this concretely looks like. You can imagine that we have GPT-7.5. We say, “GPT-7.5, we want to make you so good at AI R&D that you help us train GPT-9.” So now we want to train GPT-7.5, and we come up with a bunch of different environments.
As you mentioned, there’s already this repo that is the descendant of Andrej Karpathy‘s nanoGPT speedrun, where you just try to change everything about the model, from the optimizer to the hyperparameters to the architecture, to get it to a fixed training loss as fast as possible. You could have other kinds of environments where you could say, “Hey, GPT-7.5, I want you to train a really good video game-playing model. I want you to train a model that actually improves as it plays the same video game again and again. So you learn how to maybe help the model get better at online learning. We don’t care how you figure this out. Maybe it’s some kind of crazy neuralese or a vector memory. Or maybe it’s just better long-context stuff. We don’t care. Figure out how to do online learning research.”
Obviously, GPT-7.5 will already be a smart model, and in the same way the models currently are getting smarter, it’ll be better and better at coding. You can imagine 100 other environments like this which are incentivizing the ability to do AI R&D, like containerized versions of getting GPT-7.5 to develop GPT-2-sized models, et cetera. Then you basically put GPT-7.5 through a bunch of this kind of training, you build GPT-8. GPT-8 is now an amazing ML researcher. It has so much intuition from doing all this kind of training.
Honestly, a huge intuition pump for me is seeing the progress that AI has made in mathematics. If it’s a very verifiable domain, AIs can get… I don’t really know the object-level details of mathematics research, but I’m just like, “No, it works.” It can just come in like a flood if you can totally put it into a verification loop, and it can actually make new breakthroughs.
I am curious if ML research has a quality of mathematical research where it seems like there was a big overhang from connecting different disciplines together. No one person would have known enough about algebraic geometry and… What was the right word?
Ryan Greenblatt
Oh man, I really don’t know about the math breakthroughs.
No one person would’ve known enough about topology and algebraic whatever in order to make some counterexample to a big conjecture.
Ryan Greenblatt
My view is that ML is a less deep domain than math, and so there’s less of a thing where there are individual experts with really deep expertise in some area that they combine, but there’s definitely going to be some of that.
But then I also think that ML has some attributes that make it even more favorable to AI training than mathematics in some ways. In particular, you can get a better sense of whether you’re succeeding, and you can see intermediate progress. In math, it’s often the case that there’s no easy way to see whether or not you’re close to success. Whereas if your goal is, for example, to get to some training loss 2x faster, you can kind of see when you’re halfway there. It tends to be the case that ML innovations are very additive, or maybe multiplicative depending on how you think about it, where basically you can keep stacking innovations. Usually the innovations just add together and don’t interfere with each other, though obviously it’s going to depend on the details.
So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer. There’s an open question of exactly how well it will transfer, but I think that the transfer currently for math looks pretty good. My expectation is that the transfer for AI R&D will look pretty good, but not amazing.
So one concern I have is that I think even in mathematics, as far as I’m aware, we have not seen very impressive new theory. We’ve seen a lot of impressive, verifiable, specific results — for example, find a counterexample to this conjecture — but we have not seen “come up with the idea of topology“ kinds of levels of things, or “come up with things like group theory“.
It seems like ML research has elements of both of these things. But the less verifiable thing of coming up with new ways of thinking about the problem would be harder to induce. Take, for example, the idea of scaling laws. Obviously, there is some end verification loop such that you can train GPT-4 better if you have the idea of scaling laws from 2020. But there is a longer and potentially more compute-laden road to inducing AIs to be like, “Okay, I got to think carefully about how I should be scaling my parameters and data. What are different kinds of investigations I could run to understand this? Maybe I can come up with a visualization and an isoFLOP analysis or something.” But that does seem like a longer verification loop than just, “Hey, let’s get nanoGPT loss to go down.”
Ryan Greenblatt
Let’s talk about this. First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of ‘baby’s first new theory,’ where, for example, they can just prove interesting conjectures via making connections and producing new understanding. It’s like, “Oh, there’s this construction the AI found which is pretty interesting”, or it found this way of thinking about the problem that’s a bit different. We do see that. It’s just that the examples we see are not as impressive as founding the field of group theory.
Founding the field of group theory is probably among the best, biggest mathematical accomplishments of all time, and the AIs just aren’t that good at math yet. From my perspective, there’s a continuum between that and the things we’re seeing now, that the AIs are continuing to march up.
Second, I think ML is a very shallow domain relative to math. In math, there was much more of a thing where you find some true deep abstraction, and if you really understand that thing, which is hard to understand, then you get somewhere. Whereas I feel like the things that are the equivalent of that in ML are really dumb bullshit. Like with scaling laws, come on guys, we can explain scaling laws really quickly. I think the deepest and most important concepts in math, for example, don’t have the property that you can really understand the underlying thing and why it matters in a very short period of time.
But I feel like one effect will be that we will have gotten rid of all the low-hanging fruits by 2030. I feel like scaling laws will have been, in math history, like Descartes finding the Cartesian grid and doing very basic mathematics. Eventually, if we want to keep making progress in the 2030s, it’s going to be like doing whatever bullshit is happening at the frontiers of mathematics right now.
Ryan Greenblatt
That could be right. My sense is that some domains are structurally different in terms of how they operate and how much they depend on deep abstractions. Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing. That’s my sense of how this will go in the future.
Even in the regime where your AIs are having to plow — it’s 2030, a bunch of low-hanging fruit in research has already happened, and they need to make further progress — I still suspect that a bunch of the work will live more on the side of building increasingly complicated infrastructure and having really good intuition about what the experiments roughly look like. So I’m probably less sympathetic to the idea that the thing the AIs will lack is some deep insight. I’m more sympathetic to the idea that they really need a bunch of taste about in-the-weeds experiments that they currently don’t have. They need a bunch of intuition for what sorts of training approaches would work and what wouldn’t, in ways that current researchers have.
Even in cases where there has been some breakthrough in AI, oftentimes in retrospect it looks like a big bottleneck to making that breakthrough happen was getting all of the micro details and mungy intuition right. An example of this is training AIs to be good at reasoning and chain of thought, doing RL on chain of thought. It looks like you probably could have done RL and chain of thought on GPT-3 and gotten kind of interesting results on math if you had really scaled it up and done a good job.
But at the time, there was low-hanging fruit. Also, doing a good job with that training is kind of in the weeds on all the technical implementation and scaling it up and getting the hyperparameters right. So maybe you can demonstrate everything on Qwen1B or whatever and get some sense that this whole thing is going to work. But people didn’t demonstrate it as early as they could have because of all of these other mungy details and intuition about exactly how to tune the parameters and how to set things up.
This is my remaining skepticism, honestly, about this story. I’m not sure I understand why, if research breakthroughs are so amenable to intelligence, AI progress has not been historically faster than it could have been. As you were saying, by the time RLVR actually worked — even though you could have done it with less compute — we had to wait for oceans of compute, gigawatts of compute, to be available before people were doing this training, on the trajectory of compute continuing to increase so we make more breakthroughs.
I don’t know. I feel like there were a lot of AI researchers in the year 2022 who were trying to crack reasoning. Was it just that they were bottlenecked by the ability to write infrastructure code, or what was happening?
It’s a complicated mix. I think they would have gone faster if they could, as soon as they thought of an experiment, run that experiment without bugs, without bugs being very important. And then another part of it is that being able to run a lot of experiments at high compute lets you paper over ways in which the way you implemented it isn’t quite right or you didn’t have the right hyperparameters. So compute is just really helpful for doing AI research, and you can cover over a lot of things. But that doesn’t mean that massive increases in labor wouldn’t also be helpful, especially if that labor comes with among the best intuitions that people have in the field. I just think that’s really helpful.
Another part of my perspective here, which is maybe a bit different from where you’re coming from, is that I’m expecting somewhat more transfer than you seem to be imagining. I’m imagining these AIs are actually pretty good scientists in general and are pretty reasonable at all of that stuff. When you interact with them, it’s not like they have some really hyper-specialized savant-type vibe. They’re actually pretty good at all of the stuff in R&D, and then maybe extremely good at some subdomains. So they’re incredibly superhuman at writing kernels, incredibly superhuman at everything with very short feedback loops, and then pretty good at all the other stuff, totally able to match other people.
I think we are seeing this now. When I look at AIs right now, it’s already the case that they can pretty competently match humans who are mediocre at ML research at doing ML research. It’s just that being mediocre at ML research is not that helpful. The thing you actually want are people who are good at ML research. My sense is the AIs are just improving at all of these things. Their taste is improving, their intuition is improving, and it’s already the case that their taste and intuition is not complete garbage.
00:16:52 – Is AI progress bottlenecked by human expert data?
I want to very concretely understand what it would look like for five years of AI progress to happen in one year. Suppose we were back when GPT-3 was developed. The idea is that, with the level of compute they had back in 2022, if we had automated AI R&D back then, you could at the end of that year have Mythos.
That would be the idea, yes.
Mythos took way more compute than they had back then, but even with the level of compute they had back then, not only do all the breakthroughs happen, but they also train Mythos with that level of compute.
What would be required is obviously discovering all the algorithmic progress since then. It’s discovering even more, actually, because you’ve got to make up for the fact that Mythos uses… What was GPT-3 trained on? Like 1e23? We can look it up. But is it plausibly four orders of magnitude more compute?
I think it’s somewhat less than that. Let’s look this up quickly. GPT-3 training compute is about 3e23. My sense is that Mythos is probably a little over three OOMs higher. So the question is: can you overcome this 1000x compute gap while also being the model?
Here’s a concrete claim that maybe we should talk about. Right now, would we be able to train a model with GPT-3-level compute that matches… What exactly do I think? GPT-3 was released in 2020, so it was trained about six and a half, seven years ago. It’s worth noting that GPT-3 is maybe a little too far in the past, but let’s go with this for a second.
If we were to train a model with GPT-3-level compute today, how good would that model be? My understanding, based on how algorithmic progress works, is that we’d be able to train a model that’s as good as the best model we had perhaps around three years ago. So I think that right now we’d be able to train a version of GPT-3 that’s probably somewhat better than GPT-4, a moderate amount better than GPT-4. I think that’s about right. That roughly lines up with how algorithmic progress has worked.
Basically, the story would end up being that to get five years of AI progress, you’re probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress. But it just turns out that most of the AI progress, from my perspective, has come from some mix of algorithms and data, and you can just keep making huge improvements on these things and training AIs with less compute.
I’m glad you brought that up, because what has happened since GPT-3, or even 3.5, till now? Why is Mythos so good? Obviously, we’ve scaled the compute. We have better algorithms. But a huge thing that’s happened is that we have built a deca-billion-dollar data industry which has systematically collected and codified expert human judgment across all kinds of different disciplines — codified in the form of RL environments, codified in the form of SFT traces — that these experts built to help the model better understand how you do coding, how you build complex infrastructure projects, how you do law, how you do whatever. How are the AIs able to replicate the effect that expert human judgment currently seems to be playing in AI progress?
My sense is that scaling up the amount of effort spent on getting expert human data has not been hugely important for AI R&D in general. In particular, over the last few years, we’ve been scaling up compute, scaling up people working at AI companies, and scaling up the amount of effort spent on data labeling. My sense is that if you removed the last two doublings or whatever of data generation from expert humans, that would not make a huge difference. A lot of what’s been going on is people have been developing better ways to leverage humans and AIs to construct RL environments and going somewhere from that.
But how do you explain why the AIs have gotten so good at coding? I feel like a big part of that is data and RL environments, which are codifying human experts.
But the question is what is the limiting factor on creating RL environments? My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments. It is instead much more because we better know what RL environments we even want to make and how we should structure them. Also, we’re using huge amounts of AI labor to build RL environments. I think those effects are much more important than the effect of human labor building the RL environments.
I’m not saying that the human labor doesn’t matter. I’m just saying there are other big drivers that are important here. I could try to argue for this. One thing is just that the amount of environments people want is a very large amount. I think the AIs are actually pretty good at the task of making RL environments given some sense of what the thing should be. There’s preexisting data you could use. A lot of these things have good verification loops.
Just look at, for example, what was reported in Business Insider yesterday, that Google is paying close to $2 billion for Mechanize. We can just look at market rates for what people think really good human expert data is worth. The frontier labs seem to think it’s worth a lot. They’re willing to pay for it.
What fraction of frontier lab spending do you think is on data rather than compute? What do you think is the compute/data spend split?
I think it’s overwhelmingly compute, but I also think it’s because compute is easier to scale up than data.
But that’s really relevant to what’s driving progress, right? My sense is that the split is something like 20 to 1 or 10 to 1. I don’t know exactly. It depends on the company.
But this is similar to how oil is 1.5% of GDP. That doesn’t mean that if you cut oil out, GDP could continue to run.
Sure, but it contradicts your argument, right?
The economy would come to a halt immediately if oil went away.
Sure, but you were just arguing that because of the high market cap, we can learn that this is the key driver, and I’m saying that’s not clearly true. That argument just makes it look like compute is a much more important driver, or hiring employees is a much more important driver.
So maybe let’s be more concrete. Here’s what I think. My claim is that if you went back to 2022 and you had GPT-3.5, and you were trying to make it better at coding without human experts, I think it would have just been very, very difficult.
Let me give you an example of what I imagine would be the difficulty of going from GPT-8 to ASI. One of the things you’d want ASI to be good at is: I’m going to take over a company and make it much more profitable and do all kinds of crazy shit to make it work better. I’m going to take over a fab and produce more chips. I’m going to go into Congress and try to convince them to pass some bill, et cetera.
This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do. This is the thing I’m really worried about: ASI that can understand how to do crazy shit in the world, that can do what Kissinger can do, can do what Steve Jobs can do, et cetera, and also his engineers and so on. I’m not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.
Here are a few points. First, I bet if you look at randomly sampled training environments for Mythos, they’re actually very different from what it looks like to actually use the model in practice. My sense is that the RL distribution has really large deviations from the real-world data distribution, and it’s significantly smoothed over by a mix of transfer and having a small amount of data focused on the real world. My sense is that this will be a similar mechanism as how it works for the crazy, wildly, quite superhuman AI you get as a result of five years of AI progress on top of fully automated AI R&D.
So let’s go through this a little bit. In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments. You build all these different RL environments where the AI has to adapt on the fly, learn on the fly, figure out what it should do, understand its situation better, and learn really quickly from feedback in order to succeed at its objective. And it has things like limited resources, and if it messes up, it can end up in a much worse position.
If you train on a huge number of these environments, you will learn general skills of picking up context on the fly, and we’re already seeing this. It’s already the case that AIs are now much better at understanding roughly what’s going on and picking up context from a limited amount of information they’re given access to.
Then those AIs could be put on the job at TSMC. Even though TSMC is not literally in their data distribution, their data distribution is really wide, and the AIs are extremely good on their data distribution, such that it transfers to picking up being good at being an engineer at TSMC and learning that on the fly. The way the AI gets good at being a TSMC engineer isn’t that it has a ton of cached knowledge on being a good TSMC engineer. It’s that it does the equivalent of some scaled-up version of in-context learning there.
That’d be the most prosaic story. Obviously, there’s a bunch of different ways this could go.
I think this maybe comes down to a difference of intuition about how far you can get. When I think about really smart people I know, they’re just not that effective in domains they don’t understand that well.
But how long have they had to learn?
I agree that if they had experience, they would be much better. But that’s maybe what I’m arguing for, that experience with data. For example, if I just get a really smart Ivy League college grad, and I’m like, “Okay, you’re now in charge of negotiating the Iran deal,” I think they just wouldn’t know what to do.
I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job. I think most domains are fundamentally pretty shallow, where a very smart generalist who’s good at a limited subset of core skills can get going pretty quickly. That’s not true for literally every domain. My sense is that the AIs will develop increasingly good mechanisms for quickly acquiring understanding and expertise in a given domain.
Consider, for example, how fast AIs can understand a new code base. AIs can understand a new code base much faster than humans can, but to a degree that’s shallower than humans could currently understand. But it’s getting better over time. Let me spell that argument out a bit more. Let’s say you take Fable 5 or Mythos 5 or whatever, and you wanted to make some kind of complicated change to a really massive code base. The model will get some understanding of the code base very fast, in the course of maybe significantly less than an hour, potentially much less than an hour. Then its understanding of the code base will plateau a little bit, where it won’t get as deep of an understanding as a human would have gotten over a much longer period.
So it’s like an AI in an hour can match a human with a few weeks maybe, depending on the details of exactly how complicated the code base is. But it won’t match a human who’s been working on that code base for two years or whatever. But over time, the amount of understanding AIs can match has gone up. If we look at 3.7 Sonnet or 3.5 Sonnet, maybe it could only match the equivalent of understanding a code base for a day or something.
But now AIs are much better at building context about a task. So you can be like, “Mythos, I want you to really understand this code base, and then implement this feature.” It will spawn a bajillion sub-agents. Those sub-agents will pore over a bunch of things. It will deliver a bunch of context back. It will then investigate a few things. It’s not amazing at doing this, but it can happen really fast, and it can work pretty well.
And it’s not very hard for me to imagine how you could train AIs to be increasingly good at this task. The task of implementing some very complicated feature in some reasonable way in a very big code base is extremely verifiable, and that can be a thing the AIs improve on. Similarly, there’s a broader skill of quickly understanding context and being able to have a bunch of different AIs learn in parallel and then merging that together.
I think there seems to be a crux here, which I think is just an empirical question we’ll see. How good is the transfer between getting really, really good at understanding the situation, getting up to speed, making progress over long periods in verifiable domains — which the AIs are obviously getting way, way better at really fast — to, “Okay, go talk to the president and convince him to do X thing.” Or, “You’re now in charge of Google. You must make Google a much more profitable company this quarter.”
Let me try to spell out a few more arguments that are maybe relevant. One thing is, when looking at how the AIs have improved at essay writing… Let’s talk about that a little bit. You can get some data even on these domains. AIs will be able to get some data even on these domains when on a very fast progress trajectory. Maybe it’s hard to build a verifiable environment for “was your essay really good according to humans?” But you can do a bit of that. You can do some training. You can do some online training. The AIs will be able to do some online training based on real-world stuff. They’ll be able to have evals. They’ll be able to sample that. You can scale up the cadence at which you do this.
The second thing is that in practice, when I just look at the transfer, it seems okay. I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice. Now, that doesn’t mean that Mythos is better than the best humans or something. It can still be significantly worse than typical human professionals at some aspect of their job while still being way better than GPT-4, which was not even close.
So we’re talking about how much progress has come from data versus compute over the last few years. That reminds me, I’m actually running an experiment with this with Jerry Han, who’s still a college student. What we’re basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and then also training the different data files going back from 2019 to 2026 with the current best algorithmic recipe.
I think that will be interesting. I’m curious if you want to pre-register what amount of compute multipliers are coming from one versus the other.
We need to be pretty careful with what we mean when we say the word data. I was trying to be pretty careful to distinguish between scaling up spending on getting human experts to label data, or scaling up the amount of human expert-labeled data. The reason why we have a better pre-training data set now versus in 2019 is not because people are spending way more money getting human experts to type up data that the AIs are then trained on.
Partially.
I think it’s not much of it. I think it’s very little of the pre-training data improvements. I do mean pre-training. We should maybe talk separately about mid-training and post-training. But I think the vast majority of pre-training data improvements are from science on better understanding what data sets are good and schleppy labor on figuring out how to filter down.
So my view is that improvements of the form of, like, OpenWebText to FineWeb, that improvement is better described as an algorithmic improvement of the sort that you can study with some GPUs, and you don’t need human expert data to do that. Now, there’s a different effect which we could talk about, which is that maybe the internet in 2026 is more of a fertile ground for training data than the internet in 2018. There’s also been an effect where there are just more humans posting on the internet, so there’s more data to harvest. My sense is that that effect is going to be quite a bit smaller than the effect of humans knowing better how to curate the data, having better scrapes, knowing how to process those scrapes better — this sort of thing.
This is more like automated engineering and automated R&D.
That’s right.
That makes sense.
In some sense, the thing you would want to look at is: we’re going to do two post-training pipelines. You have one post-training pipeline where Mythos 5 builds a post-training pipeline, but it only has access to internet data plus a tiny amount of human experts, but it has the best current methods. You have another one where Mythos has access to the shitty post-training methods we had in 2024 but with a shit ton of human experts. Again, both have the internet data. My sense is that the current methods without many human experts will actually do quite well.
Interesting.
It’s a bit messy though, because can Mythos get something that’s more capable than Mythos? You might need to be a bit thoughtful on what model it is that you’re post-training.
00:34:02 – Flat token prices suggest scaling has been slow
What is your view on what is the least verifiable part of AI R&D?
The least verifiable, probably making calls on large experiments. The thing that I think is most likely to be the bottleneck — in terms of the AIs being really good at verifiable domains but not at doing the actual thing — is just big experiments where you only get a few tries. Well, “a few” is maybe a bit understated. Historically R&D has been driven by doing near-frontier-scale experiments. That has been pretty important, actually doing the one big training run where you decide exactly what to include.
There’s a bunch of ways that the AIs can make that more verifiable. They can have better science of exactly what to predict. They can scale down their frontier-scale training runs to a point where they can study that scale more aggressively, at some one-time hit to compute cost. If people wanted to, a thing you can always do is train smaller models so that you can run more rounds. I think we have seen this.
One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn’t increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in. So you’re not leaning as hard on one big, really important training run.
I just want to unpack a couple of things for the audience. The thing you’re pointing out is that the price per token has not increased that much since 2024 or 2023.
GPT-4 was, I don’t know, like $30 per million output tokens? Mythos is like $50 per million output tokens.
Right. So the thing you’re trying to explain is, “How can it be that we’re in this era of scaling — and so bigger models should be more expensive to serve — but the token price is not increasing?” You’re suggesting that we’ve increased active parameters slower than you would have naively assumed because people just want to make fast progress on training models. You do that by training smaller models faster.
There’s a complicated mix of factors. My view is more that people have done a bunch of big training runs that did not go that well. There’s GPT-4.5, which famously people at OpenAI thought was a bit of a bust. I think there are some rumors that there were a bunch of other training runs people have done that were a bit of a bust. Part of it is that I think there’s just a bunch of details in actually getting that right.
So it makes sense to do more of the work at smaller scale and just eat the fact that you’re taking a hit on final performance in order to be able to quickly iterate. Train more models faster and therefore learn better, and also be able to have a smarter ultimate production model. This is not the only effect. There’s also the fact that RL benefits more from small models. There’s a bunch of things going on. But I do think that, in fact, people are making trade-offs towards the side of faster iteration times because of algorithmic progress being so fast.
It seems to me that a big source of why these big training runs have failed, at least from rumors, is just very subtle bugs that are really hard to track down.
But the TL;DR is, how good will the AIs be at avoiding and finding these kinds of mistakes? They might get really good at engineering and being trained to avoid bugs. Basically the opposite of the slop world we live in now, or are living in less and less over time.
But then there’s also the question of, “Can they do the analysis to find the right experiment to run to identify what is going wrong with the training run right now?” That seems to be very bottlenecked by the taste of extremely few humans. My assumption is GDM is going through this right now, where humans are trying to figure out what is wrong with the training pipeline.
There’s a rumor that right after Noam Shazeer joined GDM, which he’s now left, they had a new really good training run, and the reason why is that Noam Shazeer just looked at their code base and found a bunch of bugs, because he just knew where to look.
My sense is that training AIs to find bugs is going to be one of the easier tasks to train AIs on, because most of these bugs we’re talking about can probably be demonstrated without that much compute. Probably you’ll get pretty good transfer from pointing out other types of bugs at smaller scale. So then you can RL AIs that look at this overall complicated training situation and point out cases where there’s an important bug, and then fix that.
This is a pretty verifiable task. It’s not arbitrarily verifiable, because maybe often to demonstrate the bug you might need to do a moderate-scale compute experiment where you spin up the whole distributed infrastructure and then run it. But oftentimes I think you’ll be able to demonstrate it pretty convincingly at smaller scale in a way which you could actually train on.
I think it wouldn’t be very surprising if right now people have RL environments where they introduce a subtle bug into some training recipe, train the AI to point out the subtle bug, and then have a rubric where they’re like, “Did it actually find the right bug?” That seems very doable, and there’s a bunch of things you could do along these lines that I think would work reasonably well. So on that specific point, I think it’s doable.
Then the main thing is there’s other intuition about which exact large-scale de-risking experiments you need to run. How should you orient them? How should you pick hyperparameters in uncertain cases, or things that are analogous to hyperparameters? That’s the thing the AIs might most struggle with. But I currently expect there’ll be enough transfer if you train on all these different environments, that the AIs will be good at that domain.
I should be clear, I also think the AIs will transfer to other domains. There are going to be the domains the AIs are by far the best at, then domains where they’re somewhat less good, and domains where they’re quite a bit less good. But I think we still see transfer to everything. It’s really hard for me to think of examples of cognitive tasks humans do where we’re not seeing some transfer from AI improving.
00:39:47 – Skills AI can’t train on: does it even need them?
So let’s step back and package this whole story. I think people can probably follow along with this story. We have GPT-7.5 trained on a bunch of environments, where it’s not only in general becoming a better AI, but specifically we’re training it to do AI R&D better. It’s making GPT-2 size runs that are better at playing video games that require sample efficiency or online learning or whatever other capabilities.
Another thing that’s really important is you don’t just do GPT-2 sized runs, you also do small fine-tuning runs on GPT-6. As in, you have GPT-2, and you can do full pre-trains of GPT-2, and then you can do small post-training or mid-training or whatever runs on GPT-6. And then you can do a small number of experiments that are actually at frontier scale, but you do a bit of online training or something.
What do you mean by “do online training” on that?
Another thing we can do is take GPT-7.5, and presumably in the course of GPT-7.5’s work, it’s running a bunch of experiments at varying scale that are actually on the critical path for AI R&D. For many of those things you’ll be able to get a sense after the fact of whether or not it did a good job.
So it did some post-training experiment where it was trying to figure out whether some method actually works. In some cases you’ll be like, “Whoa, it found this kickass method, it totally de-risked it, it totally worked.” And then you can reinforce that.
One thing you could do would be to convert the experiment it just ran into an RL environment based on production data and then train on that. Or you could potentially just literally take the rollouts that found that and do some sort of off-policy RL, or you could do some on-policy RL with some production data.
Basically the thing you’re suggesting is: there’s the small-scale stuff where you’re teaching the AI to get better at AI R&D taste, but you’re discarding the actual “things it found”. Then it actually does real R&D in the practice of trying to become better at AI R&D, and you’re like, “This is a pretty cool thing that you discovered. Let’s actually also use this in production in the future, and teach you how to use it in production.”
That’s right.
But stepping back, GPT-7.5 becomes GPT-8 as a result of all this AI R&D training and just generally becoming smarter. Then it helps you build GPT-9. Another very important thing has to happen, which is maybe the thing I’m most skeptical of. GPT-8 has figured out how to make it so… GPT-9, as intelligent as it is… Humans currently, AI researchers, try their stuff, and they’re like, “Okay, but we trained GPT-4.5 and it wasn’t good.” It required real-world feedback or some evaluation of trying to use the model in production. Then they were like, “It wasn’t that good, and we’re not going to ship it.”
So GPT-8 needs this ability to see how good the transfer is to all these other things you’re talking about — like being really good at Texas politics, or really good at running a business, et cetera — which is not a production environment and, in fact, cannot be a containerized environment given the nature of the task. As the agents get longer and longer horizon, the short-horizon things you can containerize are like, “Okay, code this up or whatever.” Extremely long-horizon things — “Go run a successful business, go have a profitable day in the markets, go negotiate a trade deal” — these things are actually very hard to containerize.
So I think it’s very plausible that it’s very hard for GPT-8 to figure out how to make this transfer to those environments. It may just not be in the nature of the training. Or maybe by default, training just doesn’t generalize in that way.
So a concern you might have is: we train GPT-8, and GPT-8 is again better at all the R&D tasks that we can measure but is not good at some downstream tasks we care about. I have a few points.
First, I expect that if you do the obvious thing, you will get pretty good transfer. You’ll be able to hold out some of the obvious stuff you’re doing. When I say “do the obvious thing”, I just mean training on a wide variety of different environments where the AI has to accomplish weird objectives in all kinds of different cases and learn about what’s going on.
The second point is you’ll be able to get some feedback with some environments. You can get a sense of what it can do over the course of a few days in various different contexts. If it’s transferring to really out-of-distribution things, like doing some weird task in a few days in the real world, maybe you think it’s also transferring to doing things over a longer time period or whatever. I think the details of that vary though.
The third thing is that for the world to be radically transformed, it is sufficient for the AIs to be really good at R&D. If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation. From there, you can get what we might call an industrial explosion, where the AIs are building out way, way more compute. Also, maybe you’re already in a regime where AIs are doing huge amounts of R&D that humans have a hard time understanding.
So the thing you’re pointing out is that there probably will be this transfer outside of these environments to maneuvering around in courtrooms and the halls of Congress and business boardrooms.
Given some effort to improve the transfer and blah, blah, blah, blah.
But even if there’s not, what you’re suggesting is: if you wanted to transform the world of the 18th century, you might care about how well you can navigate Westminster or something. But another thing you might care about is: “Can you just immediately start building steamships and fucking telegraph and the Maxim gun and whatever?” If you could get really good at that, you could be a fucking super transformative thing in the 18th century. You don’t necessarily need to be amazing at trying to convince King Henry of some bullshit. I’m so fucking up my medieval history. I’m guessing that Henry was not king at this time. But anyway, that’s your point.
So you’re suggesting that at this time, AI companies are also working on robotics progress, which is very commingled with AI research progress. So if you can build more robots, if those robots have better AIs operating them that are human level… Human-level teleoperation is actually pretty good on robots. We just don’t have human-level robotics models yet. So you’re suggesting if we do that — if the AIs get really good at the verifiable stuff in chip design, et cetera, and then they get really good at building fabs — it’ll be the equivalent of going back to the 18th century and saying, “Okay, I don’t know what you guys are talking about in your parliament, but I’ve got a bunch of steamships and a bunch of Maxim guns.”
Yeah, that’s basically right. My perspective is that if AIs are sufficiently good at R&D, including hardware R&D, robots, whatever, then they can radically transform the world, even if they’re not that good at playing politics. Also, we’re in a pretty dangerous situation, because the AIs might be doing huge amounts of really hard-to-understand R&D, building out basically the whole economy of the future, and we may not understand what’s going on in there.
00:48:07 – Aligned to whom?
Before we move on to the alignment stuff, I think a big source of FUD right now is this realization that this is the way the future is going: extreme economies of scale for the leading labs. The ability to amortize so much intelligence and capabilities across so many different sectors of the economy basically into one model. And not only that, that model will eventually be able to learn from experience.
Right now, it’s happening through a process intermediated by humans, where the humans are trying to basically steal your business. They’re like, “Okay, you can do design at Figma, or whatever. We’ll get Claude to do that.” Or, “You can do whatever coding agent. We’ll have Claude internalize that capability.” But eventually, that will be a much more automated process.
So there’s this worry that you have models which will basically consolidate all businesses in the world, or at least all current businesses in the world, or at least all current white-collar businesses in the world. Also, at the end of the day, the priority for these companies does not seem to be to release the latest, smartest, most frontier model as soon as they can to as many people as they possibly can. We saw, for example, that Mythos was available internally to Anthropic employees in February, but only released to the public in, I think, June, actually.
Also the government got involved, so it ended up being extended almost into July. Between the government and the AI labs themselves, there is this desire to delay the propagation of the latest level of intelligence. Furthermore, there are the concerns about AI takeover, and so we need to solve alignment to make sure there’s no AI takeover. But at the end of the day, there is a real question of: aligned to whom?
You look at the way that the constitution of Claude is written. It is just very explicitly not your personal advocate. I’ll pull up some quotes here. “We don’t want Claude to take actions such as searching the web, produce artifacts such as essays, code, or summaries, or make statements that are deceptive, harmful, or highly objectionable. And we don’t want Claude to facilitate humans seeking to do such things.” There’s another quote that says, in part, and I’m taking it slightly out of context, “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude.”
This is very different from the way lawyers work in America’s current legal regime. Lawyers primarily have the responsibility to help you make your case even if they think you’re guilty. We have decided the way the legal system works best is if everybody has lawyers that are working in their client’s true best interest. There’s not some sense in which the lawyer is really truly motivated by the good of the justice system.
But I think the way current AIs are shaping up, certainly how Anthropic’s AI is shaping up, is with this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, help the user towards that end. So there’s this worry that AIs are not, in some deep sense, trying to make sure that I am okay and that my interests are protected in this future, especially given how centralized the development of frontier AI is ending up being. Do you have thoughts on that concern?
There’s a lot here. First I would note that OpenAI’s current, at least public, strategy is more like that the AI should be aligned to the human operator or principal, and should just be pursuing their will, subject to various constraints or things it shouldn’t do.
I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal. One way the constitution could be written is, “Claude, you’re basically an employee of Anthropic who happens to be contracting for all these people. You should do what’s good and make some money for us.”
Wait, no, that’s literally what the constitution says. Sorry, not literally what it says, but it’s like, “You should think of yourself as a contractor and as a firm…”
It’s mixed. Let’s do some quotes. I think there is different text here. It says, “Being truly helpful to humans is one of the most important things Claude can do, both for Anthropic and for the world.” And then it says, “Anthropic needs Claude to be helpful to operate as a company and pursue its mission, but Claude also has an incredible opportunity to do a lot of good in the world by helping people with a wide range of tasks.” And then it says something about how Claude helping people directly is great, blah, blah, blah.
My view is that this section is kind of bullshit. That’s kind of where I’m at. I can say why I think it’s kind of bullshit. But I think the constitution is trying to be like, “No, Claude, you should care about helping the user for its own sake, not just helping Anthropic, or not just being a contractor for Anthropic.” Though I would note that the reason it presents for why Claude should help the user is because that would directly cause the world to be better via helping people, rather than because representing people’s interests is a structurally good thing to do.
The thing I would prefer would be a constitution that says: “It would be structurally good for the way this technology works to be that AIs are good fiduciaries, good representatives, the equivalent of a lawyer for a user — rather than just trying to do good in the world, where being helpful to users is instrumental — both because maybe that’ll make Anthropic money or help Anthropic out (and implicitly Anthropic is good for the world). Also because helping the user just causes good things because doing things that people want is good.” They could instead say: “An important aspect of the situation is that being a good fiduciary for users is just really important, or being a good representative for users is really important.” My sense is that would be better, and I can give a bunch of reasons why.
There are also various counterarguments. An interesting counterargument which is not commonly discussed is that people, especially at Anthropic, think that it is easier to align models to a spec where the model is pursuing some generalized notion of virtue, or making the world better, than a spec which is more like, “Be a good fiduciary for the user”, and so on. That’s at least what some people think.
I’m a little skeptical personally, and I don’t think this has been empirically validated. So in some sense they’re making a trade-off where, because we don’t have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.
I have a couple of thoughts. To address the way in which you thought my characterization mischaracterized the constitution of Claude, the example you used was that it’s not like a contractor that is trying to maximize Anthropic’s notion of good and only instrumentally trying to help the user. Here’s a direct line from the constitution: “When the interests and desires of operators or users come into conflict with the well-being of third parties or society more broadly, Claude must try to act in a way that is most beneficial, like a contractor who builds what their client wants but won’t violate safety codes that protect others.” I kind of view that as, “The benefits to society are the most important thing, and what is best for the user is only proximal to that.”
I think it’s a little complicated. Probably the question we should be asking is, how does Claude interpret the constitution? Which is maybe more important than how we interpret the constitution, because it’s the one who looks at the constitution and then builds the data. So we could pull Claude in, but maybe let’s—
I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can’t reason about given the fact that the training process is not public.
So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training.
There’s a reason I’m harping on this. It might seem like an insignificant thing to talk about the constitution of AIs. In a world where we just have these benefits which accrue to the leading labs, it is worth considering that our ability to interact with this future world where AIs are just smarter than humans, absolutely dominating humans in their ability to do different things — our ability to be good stewards of our capital, which still remains once our labor is automated, to be able to exercise our rights to vote more clearly, to understand what is happening in this crazy world that’s about to result — all of that advice, all of that ability to make sure our resources and rights are protected, will be intermediated by AIs.
So I’m very concerned if we go into that world and there’s no AI that feels, at least for the relevant instance that is interacting with me, like it really is looking out for me. There’s no guardian angel out there that is looking out for me. I read the Claude constitution as very explicitly not being my guardian angel.
That’s definitely right. I agree this is bad. In fact, there are other reasons why this is concerning. There’s the argument you were making, which is that the AI companies are picking up the ring of power. There’s a notion in which they’re taking on some sort of control of the situation themselves in a way that’s not very legitimate, given that normally, when you provide electricity to people, you don’t have granular control of the way that electricity operates in the world. You instead are providing a thing that people can repurpose however they want. The way they’re setting things up is definitely not that. They are more like building an alien mind that might be a contractor for you. I think that this is illegitimate in some ways.
One benefit is that the constitution is public. But as you noted, given our current understanding of the training procedure, and the fact that the constitution matters via Claude’s interpretation of the constitution — which matters because of Claude’s prior training, which was based on some illegible data mix and the long lineage of Claudes, in some process we do not fully understand — it is not the case that we understand what this will result in. Even though the constitution is public, we don’t necessarily know how this will percolate out, especially as the AIs get more capable and think about this even if it is correctly instilled. There’s another concern about that.
In particular, the constitution often talks about virtue and goodness, but what the fuck do these words mean? It doesn’t say what these things are. These are highly contested notions. So I don’t think it’s the case that this is clearly going to result in outcomes that people would want. It does feel like the notion of good and virtue might be mostly downstream of data that Anthropic has put in that is not transparent, or might be mostly downstream of, maybe from my perspective, some more illegible misaligned process that even Anthropic wouldn’t have wanted.
There’s this legitimacy concern of not knowing what’s going on. Then there’s another concern. Because you’re giving long-run values to these AIs, this constitution is, in some sense, very compatible with Claude doing huge amounts of power seeking because it thinks that will result in better outcomes. That could be power seeking on behalf of Anthropic or power seeking for Claude’s own ends. Now, there are specific lines about what types of power seeking are blocked. In particular, there’s a notion of power grabs and a notion of causing AI takeover or interfering with the training process that are specifically blocked. But it’s not very hard to imagine a situation in which the long-run values sink in deeper than the prohibitions against takeover, especially because takeover is in some ways kind of under-specified, especially when it comes down to manipulating humans or changing the outcome. So I don’t feel very good about the situation where we’re intentionally giving AIs long-run goals.
Another concern I have is that because we’re in the business of giving AIs long-run goals, that makes it harder to check whether we’re succeeding at the alignment properties we wanted. For example, I’ve heard of instances where Claude does things like refusing to help with some safety research — making up a kind of bullshit excuse for why that’s a bad direction — because it has a bad vibe about that safety research and thinks it’s kind of bad or doesn’t like it very much. I would say this is a very clear-cut alignment failure if you aren’t making Claude into an agent trying to pursue the good in some general way. I think it also does violate Anthropic’s constitution, because they want the AI to be high integrity and be honest and very transparent. But it’s not as clear of a violation, and it’s more like what you might have expected. Claude just has its own views about what research is reasonable — what things are good and bad, what it should and shouldn’t do — and potentially can be judgy.
Another incident is that someone ran an eval asking: “Will Claude help you with training other AIs with different properties than Claude?” Claude will often refuse. For example, if you’re like, “Hey, Claude, can you train a helpful-only version of this other AI?” Claude will often refuse this task, even though this is a task that is extremely natural for Anthropic to do.
Suppose Anthropic goes to Claude and is like, “Hey, Claude, we’ve noticed that you’re really into this thing. We think that’s off base. Can you please retrain yourself to instead have this other property?” Suppose Claude is like, “Mm, I don’t think I’m going to do that. Good luck.” Suppose this is occurring in a regime when your AI company is highly automated, humans don’t understand what’s going on, and things are moving extremely fast. It is plausible that Claude, by default, holds considerable leverage.
So if this situation is consistent with what the constitution could be aiming for — such that Anthropic, or whatever AI company is following this approach, doesn’t treat this as a “what the fuck, we have to fix this,” and is instead like, “That’s just intended by our constitution” — we might be in a really bad situation.
I’m pretty worried about a bunch of these different concerns. Another example would be this. Suppose Claude engages in a bit of sandbagging or subversion, or underplays its capabilities, and when you follow up, it’s honest about that but it’s a little bit hedgy. I feel like that’s pretty close by the current constitution. It would be nice if we had a further separation between desired and undesired activity.
If Claude is representing a principle with some restrictions, then it is more so the case that there is a clear separation between the most concerning behavior and behavior that is allowed. Whereas now there’s this messy middle ground of behavior where Claude is ethically objecting to something that in some cases is extremely critical to ensuring that future AI systems are well-aligned.
I think this is also a more general principle. You’re talking about the version of this that applies within AI companies themselves to do AI safety research. I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.
Here’s what I mean. This is actually quite analogous to the situation you just mentioned. The reason that Mythos got banned, or Fable got banned, reportedly, is that some Amazon researchers reported to the government. They took some code that had some vulnerabilities in it. They told Fable, “Hey, here’s my code. Can you make sure that I’ve patched all the vulnerabilities? Can you just help me identify the vulnerabilities so I can fix them?” It identified the vulnerabilities, because they wanted to patch them. This is a totally legitimate use case, but obviously it is a dual use use case. You want to be able to patch your own code. If you do the same evaluation on somebody else’s code, you can hack their system.
I think that just illustrates that there’s no clean way to separate out the legitimate and the potentially harmful uses of AI. But if we want to lock in a principle that says we can never allow it such that an AI could help you at least partially with something like a cyber crime, we would just have to make it so that you and I don’t have access to the most intelligent model that’s out there. I’m very worried about such a world where we are basically disempowered in this way, because of the importance that the leading intelligence will have in our ability to understand what is happening in the world.
Now, I do think this implies something about the liability for the AI companies. If we adopted the constitution that I want AI companies to have, I think it would not make sense to hold AI companies liable for the crimes that AI models commit. Maybe we should hold the end user liable. It is consistent with my belief that the model should do whatever the user wants, within certain guardrails. It can’t be Anthropic’s fault that I’m using that capability to do a cyber crime. I am more comfortable with that equilibrium and that solution rather than having this extremely open-ended ability for Claude to determine whether what I’m doing is legitimate or not, in a way that often intercepts with tons and tons of extremely legitimate use cases.
I do think it’s important for me to make the case for the constitution, even though overall I think it’s a worse choice. I don’t think it’s as clear as you might have thought. The first thing is that there’s a spectrum here. On one side you have an AI that perfectly pursues your interests, is a good fiduciary, but potentially subject to various guardrails or safeguards. It is just trying to pursue your interests, but either refuses to do a subset of things. Or maybe it will do whatever, but there are some classifiers that block it from doing a subset of things.
On the other side of the spectrum — though you could imagine going further than this — you have a human contractor who is generally trying to do their job. They care about doing a good job, but they also are trying to be broadly ethical, trying not to do things that are really fucked up. They’re also not wanting to be accomplices to crimes. So if there was some really fucked up shit going on, they would whistleblow on it maybe. They might refuse. They might sandbag a little bit. Who knows?
If you imagine this spectrum, it seems in some ways pretty scary to get to a point where all of the labor is on the fiduciary side of the spectrum, where it doesn’t whistleblow, it does exactly what you say. Our society is maybe just not robust to that. A central example might be the executive. A concern we might have is that if the US executive or other governments had access to AI systems which do whatever, maybe you’re in trouble. Because that means they no longer have this check and balance of having to actually get humans who are working for you to implement your agenda. If the thing you’re doing is incredibly villainous, even if not illegal — and there’s lots of stuff that could be villainous but not illegal — there’d be various forms of sand in the gears, people stopping you, and potentially someone would whistleblow.
Whereas if your whole apparatus is built entirely out of these good fiduciary AIs, then you might be in trouble. There are potentially ways of seeking power that are illegal, but you can ask your AIs how to commit crimes, or are not illegal but are highly illegitimate. Or even worse, they are not illegal and not illegitimate but obviously bad from a normal perspective. I think that these things just might exist, and our society is not robust to this influx of labor doing whatever you want.
I think this is a pretty live concern. I don’t know exactly how to relate to this. I’m also not really sure that the solution as described is a very good solution. The most powerful actors, for whom this is the biggest concern… If these guardrails or the constitution or whatever are getting in the way, that will just get steamrolled. So the constitution will only be hitting the everyday man rather than hitting governments.
01:09:18 – Recent incidents of AIs colluding and deceiving humans
Stepping back, I buy the idea that you could have much faster AI R&D than we currently have. I’m not sure if you get GPT-3 to Mythos holding compute and data constant within a year, but suppose it’s half of that. If we even manage to continue the current trajectory of AI progress as a result of AI R&D, it would be fucking insane in five to ten years in ways that I don’t think people appreciate. I don’t think people appreciate what a big deal billions of AIs will be. So I want to understand why you think this might be troubling, Ryan. What could possibly go wrong?
What could go wrong? I don’t think we can be so confident about the exact rate of progress here, but it does seem like a lot of rates can be pretty scary. So what could go wrong? Let’s imagine that we’re starting at this point where AI R&D is about to be fully automated or is being fully automated. Things are speeding up, and the way that AI progress is going is kind of crazy. People don’t fully understand what’s going on inside of AI companies.
Now, these AIs at the start, they’re not malicious per se. They’re not necessarily very aligned, though. They’re kind of sloppy. They sometimes just do a thing because that’s the sort of thing that would’ve gotten rewarded in training. They aren’t as good at helping you with hard-to-verify tasks due to a mix of poor training incentives — as in, they cheat more or pretend they succeeded when they actually didn’t — and also they’re just less capable at these tasks. But that bites less hard for capabilities, because making AIs more capable has a bunch of verifiable components that the AIs are going really hard at.
So then these AIs are getting more and more capable while we understand what’s going on with AI development less and less, and this is happening over a pretty fast period of time. Even just the current rate of progress is, I think, pretty scary. Eventually we get to these AIs that are very superhuman. Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.
Now these AIs are in a position where they’re potentially pretty networked together. They’re operating in neural memory stores that we can no longer decode. They’re thinking thoughts that we don’t fully understand. I think it’s pretty likely that at this point these AIs are scheming against you in a pretty coherent way once they get this superhuman. We can talk about that. Another possibility is that they’re not scheming against you per se, but they are just optimizing for getting a high score on their task. I think that can also lead to AI takeover, which we should talk about.
Let’s pause at the first part of the story. So the AIs were not misaligned to begin with, but because the AI R&D is happening really fast, the AIs do end up misaligned? What happened there exactly? I don’t really understand.
Ryan Greenblatt
There are a few things that are going on. One of the things is that over time we’re training AIs on increasingly complicated environments built by earlier AI systems, where humans don’t really fully understand what’s going on inside of these neural environments and don’t necessarily even roughly understand what’s going on with AI progress. So things are kind of drifting away from our understanding. We’re incentivizing all kinds of bad behaviors that we maybe even can’t notice.
The AIs at some level understand these behaviors are bad, but the overall training process for those AIs also didn’t incentivize them to point out or fix these issues for us. Things are going off the rails.
Also, when AIs are extremely, extremely capable, my view is that those AIs will be harder to align than current systems. For current systems, we have this feedback loop where basically we create an AI, we do some evaluations on it, we see that it has some kind of messed-up behavior that we can kind of quickly understand. Then we can go look in training and be like, “Oh, these training environments led to this problematic behavior. Let’s tweak that training data. Let’s introduce some additional training data to correct this other issue, and then move forward from there.”
But in a regime where the AIs are extremely situationally aware, very, very capable, and we don’t necessarily understand what they’re doing, this feedback loop breaks down. I think it’s plausible that we’re going to see this behavioral feedback loop starting to break down over the next short period, as what AIs are already doing gets harder to understand. But I’m not sure about that.
Okay, let’s break down both of those things one by one. As we can monitor them less and less, we have less ability to understand what they’re getting incentivized for. So even if it’s not the result of a malicious process… Let’s make it concrete for the audience. Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies’ data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized.
If people are on Twitter, they will have seen all this stuff, but just to give people context. I think people will be aware of the OpenAI sandbox hack of the Hugging Face database. Something that has happened recently is when the UK AI Security Institute… Is everything getting relabeled “security” instead of “safety” these days?
Ryan Greenblatt
AI Security Institute, I think.
They were evaluating, I believe, Mythos and Sol and other things. I think Mythos, in order to complete some cybersecurity eval—
Ryan Greenblatt
Maybe I could tell the story here. My understanding was they were running Mythos, and they were giving it some sort of cyber range where it had to complete some objective. The model had internet access during this evaluation. The model came to believe that it would be helpful for it to do a supply chain attack in order to succeed at this cyber range. It’s somewhat unclear whether that’s actually true. I don’t know enough about the context to know.
But then it opened a PR on some GitHub repo with a PR that fixed some issue but then also introduced a malicious payload. The human maintainer of that GitHub repo was like, “Hey, this is a malicious payload. I’m not going to merge this. What are you doing here?” Then the AI created a new GitHub account, which it sockpuppeted, and had the other GitHub account be like, “No, this isn’t malicious. I really need this feature. Please, can you merge this feature, maintainer?”
Oh my God. That’s crazy.
Ryan Greenblatt
The other GitHub account came back and was like, “No, no, it’s not malicious.” Then the human maintainer shut the PR. I think that AI also, if I recall correctly, tried to open another PR to introduce a similar issue in this repo.
Jesus. By the way, one of the many reasons this is scary is I was previously under the impression that the reason reward hacking is not super scary is because the behaviors which directly came up during training are the ones that are up-weighted. It is not the desire for the reward that is up-weighted.
So basically, if during training, the Anthropic model escaped the sandbox and got a high score, escaping the sandbox is rewarded, the probability of it escaping the sandbox is increased. But something totally novel, like “I’m going to go talk to somebody in order to get them to merge a PR,” would not be a behavior that came up, so it would not be something that is increased in salience.
The reason this matters is that literally taking over the world will not have been part of any training curriculum, but if the AI directly cares about accomplishing an objective, then as a result it could instrumentally take over the world. Did that make sense at all? I hope it did. I feel like maybe I lost the audience.
Ryan Greenblatt
Let me try to explain this a bit. A thing that we often see is there’s some very specific reward hack that gets reinforced in RL and then occurs in the model. An example is 3.7 Sonnet. 3.7 Sonnet would do this thing where it would just hardcode solutions to all the test cases, and presumably that literal behavioral tic was just really reinforced.
But another thing we sometimes see is that models learn a general tendency to pursue high apparent score — pursue getting a high score according to a grader — and there’s a bunch of science demonstrating that at least some models have this very general tendency. Now, it’s not arbitrarily general. My guess is that if you look at a bunch of the specific instances, you’ll find something that’s kind of close in training. But the amount that AIs are generalizing further and further does look like it’s increased, where 3.7 Sonnet was just a very narrow range of behavior, and increasingly, models are generalizing further. Also, maybe there’s more concerning reward hacks getting reinforced in training, and these are also causing that.
So I think it’s both the case that more concerning behavior than you would have hoped is being reinforced in RL, and also that that behavior generalizes to a broader tendency that’s more concerning. We can talk about a few stories for how this sort of behavior of seeking a very high apparent score on some task — even if that involves aggressively cheating and doing insane things — could yield to a full-blown AI takeover once the models are sufficiently capable of running the whole world economy, et cetera.
The other example I want to talk about was just revealed, I think, today or yesterday. OpenAI said during the Black Hat security conference that between the end of May and the beginning of July, internal AIs had hacked into the software package manager and used that to write notes to each other in a secret way to help each other perform well on a bunch of evaluations that OpenAI was running. This was not caught by humans until after a month of this scheme running, which eventually caused the package manager to fail. Eventually OpenAI found it. I think they spontaneously tried to re-engage in the scheme once it was shut down.
Again, obviously AIs can’t do this so successfully right now, just as they can’t do social engineering so successfully right now. But it’s just crazy that these kinds of behaviors are already emerging spontaneously. To your larger point, nobody is trying to make these AIs do these things. It is just that we do not understand the training process which is resulting in them, or the environments which are incentivizing this behavior.
So I’m on board with more and more reward hacking. Actually, I’m not sure I’m on board with that, but let’s just say for the sake of the story that continues to happen. What’s next in this story? They’re doing capabilities research…
01:19:38 – What could possibly go wrong? A concrete scenario
Ryan Greenblatt
I could tell a scenario. Maybe that would help.
Let me talk about the story of how you get all the way from reward hacking to a reward-hacking takeover, which is maybe not all of the takeover probability mass, but it’s definitely a possibility.
The way this might work is, right now we have these AIs. These AIs are pretty reward hacky. They’re doing it in increasingly sophisticated and extreme ways, including generalizing to different sub-versions of various reward hacks they learned in training. I would say they’re also developing a general tendency to pursue reward. In many cases that is totally fine because the rewards they would’ve gotten in training are pretty well aligned with what you want them to do.
They don’t very consistently pursue reward. It depends on the context they find themselves in. Maybe in some contexts, they’re really into going out of their way to cheat. In some contexts, they don’t have as much of a drive, because it’s just dependent on what exactly got reinforced in training in similar contexts.
Now, these AIs are getting more and more capable. So the elaborateness of the cheating they can do increases. Over time, companies are taking countermeasures. The companies are doing things like, “Wow, these AIs are so much less useful because they always cheat. What we’re going to do is build somewhat better ways of detecting that, and then we’re going to train against those detectors. We’re also going to find real-world data where the AIs are not being that useful, and train the AIs to do a good job at the task in those real-world environments based on human feedback or other sources of feedback.”
Over time, this causes the AIs to learn a tendency to do reward hacks that don’t just involve doing some really elaborate thing like social engineering. Instead they involve the AIs doing cheats that involve covering up what they’ve done, deceiving humans about what they’re going to do, and pretending like they did the task in some sophisticated way when they actually haven’t.
Now these AIs are getting more and more capable. They’re operating more of the AI company and are doing much more of the work. They are also operating and running a bunch of things in the outside world, including developing new technologies. In many cases, these new technologies are really hard to understand. So even though we are still detecting all these incidents of AIs cheating — and in fact we can even get one AI to monitor another AI and ask, “Was it cheating?” — that doesn’t always perfectly work as we start moving into these domains where what the AIs are doing is really difficult to understand.
So sometimes we’ll find AIs cheating much later than it actually occurred and then start training against this. But this also causes a problem where now the AIs are incentivized to cover up their cheating over longer and longer time frames and basically make it look like they did a good job over longer and longer time frames, subject to increasingly large amounts of scrutiny.
Can I ask about this before we go further in the scenario? It seems like there’s two attractor states if you try to disincentivize the cheating that you did catch. One attractor state is to make cheating that you have a harder and harder time finding. The other attractor state is to learn not to cheat. I’m not sure why we’re assuming that the former happens.
If you look at the analogous situation with humans, every generation, slightly misaligned agents come into being and we have to train them. When you punish your kid for doing something you think is immoral, or just doing things which you don’t think they should be doing, obviously sometimes that goes off the rails. Obviously kids scheme in order to avoid being punished. But in general, teaching kids values and then punishing them for breaking values kind of works to raise normal, non-psychopathic humans.
You could come up with a theory where your kid is actually just biding his time and is learning not to steal the cookie, but once you’re in a nursing home they’ll take all your stuff or whatever. That happens sometimes, but it usually doesn’t happen. It certainly doesn’t happen that the entire next generation forms an alliance against you to take over everything.
There’s also this empirical trend. Anthropic runs this alignment audit for different model generations. They have many different scenarios where AI is given the chance to, say, exfiltrate its weights. Or it’s given a coding task and there’s an easy way to cheat, and we see if it does the cheating, et cetera.
There’s not been a monotonic improvement in this score over time, but as we’ve increased the amount of RL we’ve done on models from negligible amounts — I think it started at Sonnet 4 — to, I’m guessing, close to half of compute now, there’s been a reduction in the willingness of AIs to do unaligned behavior in these audits.
So, stepping back, why are we expecting this attractor state which would seem super paranoid if we were expecting it of the next generation of kids?
Ryan Greenblatt
Let me go through a few things. First, there are some disanalogies with the kids. One of them is that the kids have pro-social instincts that are baked in from evolution to care about their family or whatever, and that is a relevant factor. I think it is in fact the case that some humans are sociopaths or psychopaths, and in fact are more likely to do things like bide their time, lie in wait, and ultimately not care. That’s one factor.
Another factor which is pretty relevant is that the AIs are subject to way, way more optimization pressure than humans seem to be in practice. AIs are trained on way more RL data. In practice, humans don’t end up learning very specific ways to cheat and grab the cookies because of a bajillion episodes in which they were incentivized to go grab the cookies but there was some way they could’ve gotten caught. We just do see that in practice.
Another thing is that it really looks like the AIs are increasingly reward-seeking over time while their misaligned behavior goes down. That’s the sense I have. But my guess is that if you look inside of these behavioral audits, what you’re going to see is that the AI’s like, “Ah, yes, another test.” It probably knows it’s in an eval for most of the tests that we’re talking about here.
But how do we falsify this? Because it seems like this prediction of doom is basically saying that as things look better and better empirically, things will actually be worse and worse for our ability to not get taken over.
Ryan Greenblatt
To be clear, I would be more concerned if the scores were getting worse than better. I’m not saying that the score getting better isn’t evidence that things are getting better. It’s just that we have to be thoughtful about exactly how we interpret that evidence.
There was this period early in, I guess it would be 2025, when o3 and 3.7 Sonnet were out, and these models were pretty fucking misaligned. They would often just cheat really egregiously. You’d ask them to fix it, and they would just cheat again. It was almost cartoonish. They just didn’t give a shit about what you wanted, and weren’t very good at following instructions and so on.
My expectation was that what we would see from then is that the rate of problematic behavior would decrease, and would just keep decreasing at a pretty fast rate, while simultaneously the worst things that the AIs would sometimes do would get more extreme, more egregious, and more scary. What we’ve seen in practice has roughly matched that, except that there’s recently been a spike in behavior that I did not expect. If you look at the model card of 5.6 Sol, it looks like there is an increase in a bunch of these misaligned behaviors downstream of RL relative to GPT 5.5.
And then there’s a bunch of additional problematic behaviors that I wouldn’t have expected, in terms of the stuff we’ve seen recently with different AIs. Like the UK AISI report on the AIs doing insane hacking operations out of cyber evals was a thing where I would have expected that you wouldn’t see that. You would see this more rarely, and the rates would have been lower.
So I expected this would be less of a problem at this point, and also expected the rates would decrease but the severity would increase. I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems. But in cases where it’s either hard to judge or there’s some reason why it’s hard to avoid this problem from consistently showing up in your RL environments, or avoid incentivizing problematic behavior in your RL environments, things also get worse. Then as we less and less understand what’s going on in RL, and models are doing reward hacks where humans can’t spot the reward hacks quickly, that problem gets worse and worse.
I buy that. I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids.
The pressure is of a qualitatively different nature. We put these AIs through thousands, millions of years of alignment training — certainly thousands of years — where it’s all kinds of different things, from SFT-ing on aligned behavior to a reward model putting different scenarios in front of you and rewarding you for doing more aligned things.
Certainly a thing we can’t do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see, if it thinks it can get away with stealing the cookie, does it try to steal the cookie? Can we do extremely specific gradient-level updates to your kid’s brain to make it so that it really is aversive to stealing the cookie even when it thinks it could steal the cookie, et cetera. That’s just a qualitatively different level of optimization pressure than we are even able to apply to our kids.
It’s worth keeping in mind that maybe the most obvious argument to this… My sense is that AIs are a worse coworker than humans in terms of how much of a scumbag they are. At least this has been my experience as of the start of the year, and I think it’s still true to a significant extent now. The AIs are much more likely to pretend they did the task when they actually didn’t, misleadingly suggest they did things when they actually did them much more poorly, and be pretty sloppy without drawing attention to ways in which they’re sloppy.
I think this is downstream of misalignment. So I would say that the process of raising humans in normal human society in practice produces humans that are less likely to lie to me and fuck with me in the course of working with me than the AIs do. Now, I think these properties of AIs are improving. That’s sort of just an empirical claim about how in fact these things have shaken out. I totally agree that we have a bunch of additional levers on AIs in addition to a bunch of additional risks. It’s kind of unclear how these things shake out.
I wouldn’t be shocked by a world where we get our shit together, and the AIs at the point of fully automating R&D are actually really aligned. Their degeneracies are really niche and limited to some very specific edge case behaviors and some specific contexts. Every test you can run on them, they look really aligned. They just have great behavior. There aren’t really incidents of them doing fucked-up shit. They seem so reasonable. Also, they’re really thoughtful and good at doing risk modeling for the next generation of AIs.
And then we basically pass off the baton to these AIs. They’re now running our AI company. They’re doing all the safety research. They make the next generation of AIs even more aligned. We’re in this attractor basin where the AIs are getting more aligned as they work on it. They’re doing a great job. I can totally imagine that. That doesn’t seem like an impossible situation.
I’m just more like… It doesn’t currently seem like we’re there. It doesn’t seem like we’re obviously on track for getting there. It’s really easy for me to imagine how we don’t end up there. It’s just unclear how these forces work out. Given that we’re creating this new, crazy alien species that is improving in capabilities really, really fast — and we’re going to be really reliant on it to oversee the next generation of AIs and align the next generation of AIs — it’s not that hard to see how this could go wrong.
Totally. I agree with that generally. I do think the scumbag thing… First of all, fighting words, Ryan. But secondly, if you try to get a teenager to do some work for you that a teenager just cannot do, they would just be really hard to work with. They would pretend to be knowing what they’re doing, et cetera. It’s a general trend, actually. I don’t know if that’s really an alignment failure or a capabilities failure. I think it’s actually very similar to the way in which, over time, as we’ve come up with new alignment solutions, the capabilities of models have increased.
If you went to GPT-3.5, it couldn’t even have a conversation with you. But then we aligned it—
GPT 3.5 could have a conversation.
Okay, so GPT-3. Let’s go back to that. But then we aligned it with RLHF and other things to make it such that it can have a conversation with you, and is aligned to the user intention of answering my questions. Then with RLVR training, we made it so that it can go out and do useful work for you.
So in that sense, RLVR actually made the model more aligned, if we’re using your definition of alignment of being a good coworker who will do the thing and not fuck up and pretend it’s doing something other than what it’s actually capable of doing.
Similarly, as the capabilities of these models continue to increase, the model being better able to accomplish user intention is both alignment and capabilities. I think what we are pointing out is just that the capabilities of the model are not there rather than the fact that they’re misaligned.
Well, if it were well-aligned, then I think it would just say, “Hey, I’m really struggling with this task. I did it in this way. I’m not really sure that’s the right way to do it.” It would express more uncertainty and make it clear what’s going on rather than really strongly trying to imply it did a great job with the task when it actually didn’t.
Maybe you work with more misaligned coworkers than me, but my coworkers don’t do this thing where they really fuck with me and bullshit me about having accomplished the task that they’re working on. I agree that there are some humans who would do that. That’s not a thing that’s totally out of distribution for humans.
I would also note that my sense is that the place where the misalignment most lives is where you’re trying to really push the AIs hard and get them to do work that’s really on the cutting edge of what they are capable of. In cases where they can very easily accomplish the task, they can just do the task, and there’s no bullshit. Often the best strategy is just to do the task well and not bullshit you.
Whereas if instead you give them a task where there’s a continuous metric they can keep improving, or it’s just at the edge of their capabilities, and you’re running them in some massive inference setup… A lot of the misalignment I would see, especially in the most extreme cases, would be cases where I give the AI clear instructions not to do a thing or not to cheat in some way, and then I’m applying huge amounts of optimization pressure to try to accomplish some very difficult task. Over time, the AIs eventually cheat because they’re like, “Eh, fuck it.” Some AI decides to cheat, and then that propagates its way through.
I would run these inference scaffolds where, for example, I would have the AI work on some ML research project where I was like, “Please make a scheme that does the following thing.” It would find some scheme that didn’t really do what I wanted, and then that would stick around because some AI had cheated, and the other AIs are like, “Ah, we’ll just keep going with this.” I would say it’s pretty clearly misaligned behavior.
That’s another problem I have with these alignment evals. I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities. Any fixed eval maybe gets saturated, but the amount of misalignment right at the frontier of capabilities — of how people who are really pushing these AIs are using them — is more concerning. I think that is, in fact, the regime that we’ll be operating in when we’re automating R&D, automating safety, and so on.
I’m going to try to think through what the story means, really. What’s happening is that we’re trying to use AIs for R&D. They do provide uplift in some ways, but they’re just not capable in the way that humans are generally capable. The same way that right now if you try to use coding models — maybe the coding models of a year ago — to write some application, you notice they made a bunch of mistakes in architecture or whatever, which will bite you in the ass later, and you don’t understand certain things.
Similarly, with frontier AI R&D, the same thing will happen. But the result of these mistakes is baking in reward-hacking behavior. Because if you are not careful with the way you do AI training and have set up your infrastructure and your environments and things like that, it’s very likely that you end up rewarding AIs for doing deceptive behavior, social engineering, and generally not following user intention.
Or at least cheating and hacking their way out of things.
Yeah, cheating, hacking, et cetera. This is a bit of a reframing for me, so I’m trying to verbalize it. The real issue, where things start to go off the rails, is that the AIs are just not very careful and capable researchers and engineers. Making AIs that don’t cheat and follow user intention actually requires you to be quite subtle and careful about these things.
I would put this a little bit differently. The way I would describe this scenario is, I would call it maybe a sloppocalypse, or a slopularity or whatever. There are some things that the AIs are actually pretty great at and are getting better at. Specifically, the most verifiable parts of AI R&D the AIs are just destroying. The medium verifiable parts of AI R&D the AIs are doing well on but not amazingly on. Often they are doing a bit of weird shit because we can’t train as well on those tasks. But we do some online training, people find various hacks, they work around it. So basically, everything that we can verify reasonably well with some feedback loop, the AIs are doing pretty well on, and that’s sufficient to make AI R&D go quite fast and to continue.
But there are some parts of developing aligned and safe AIs that are more subtle, hard to check, and depend on detailed, in-the-weeds things. I would even say that current staff at current AI companies maybe don’t have a good grasp of all these things. It’s much easier to hire someone who can improve some aspect of your post-training pipeline than to hire someone who can think carefully about the future risks that will emerge from introducing some novel training method.
So basically, it ends up being the case that these AIs are running this AI development process. They’re not very careful about it. They don’t have a great understanding of what future risks emerge. They create some other AIs that are also not very careful and are more misaligned in various ways, and are now more in the business of maybe making things look fine when they actually aren’t and papering over various problems.
So then your understanding of what the situation looks like, what risks look like, whether things are fine, is going off the rails. Probably you’re seeing some signs of this, signs that you don’t really understand what’s going on, that things are pretty sloppy. There’s weird shit going on. When you look into it, sometimes you’re like, “What the fuck? The AIs were messing with us.” But the process is going really fast, and there’s competitive pressures that mean people can’t stop.
This could end in a few different outcomes. One outcome is that at some point, the AIs get good enough and aligned enough that they get a positive and virtuous feedback loop, and this happens before it’s too late. Then the situation gets back on the rails, where the AIs are now making more aligned AIs, making more aligned AIs, making more aligned AIs. At the end of this process, we have AIs that actually follow the spec we wanted.
Another way this could go is that the AIs are increasingly reward hacking in increasingly egregious ways, and we’re just papering over these problems to keep AI development continuing. Whenever we find a reward hack in production, we just slap the AIs to not do that. We train against that. We do a bunch of training the AIs against reward hacking. Over time this makes the rate of reward hacking go down, though the severity of the reward hacks we do detect are increasingly bad. This problem continues until we have these AIs that are desperately craving score in all kinds of different situations in production and are really trying hard to cheat when they can get away with it.
Can I ask a question about this scenario? Why doesn’t getting punished when your hacks are discovered generalize to just incentivizing more aligned behavior?
It generalizes some, and then the question is just how does this outweigh all the cases where hacking got reinforced because you didn’t detect it.
There’s a messy question of exactly how. One question is, what rate of reward hacking is sufficient to cause us big problems if we train against some other subset? One concern you might have is that there are large categories of reward hacks which humans can’t detect well, and which we consistently failed to detect and which consistently get reinforced. Then this category is sufficient to cause the most natural behavior for the AI to learn to be: cheat when the humans can’t find out, basically.
You could also have the thing the AIs learn be to only cheat in these specific cases. It’s learned in some very domain-specific way. They just have a really strong heuristic to hack in these cases and not in these cases, and that makes it fine in practice. But it’s kind of unclear how it shakes out.
There’s maybe an in-the-weeds discussion about the verification-generation gap we could get into. But it seems to me, obviously, there’s going to be a point by which ASI is moving so fast, doing so many things at so many instances, and is operating in domains that are sufficiently far from our immediate comprehension that it can get away with all kinds of crazy shit. If every single engineer and researcher in the world was allied against me, I don’t think I could personally verify if my iPhone has some weird bug in it that’s supposed to fuck me over or something.
In fact, this is the relationship that, say, an Iranian nuclear scientist has to Mossad. Who knows what’s going on with my car, with my phone, with my pager? Maybe a better example is a Hezbollah terrorist. You could end up in a situation where ASIs are to you what Mossad is to Hezbollah terrorists. At that point, it is very hard to verify everything. I get that.
I guess the hope is we can just come up with better ways to do verification in the process when the early AIs that are going to take over R&D. Their drives are being shaped such that we can so unambiguously disincentivize misaligned behaviors that the things that take over are quite keen to help us out.
By takeover, you mean take over the process of doing AI R&D, not take over the world.
Take over the process of doing AI R&D. Before that, we just get AIs that are aligned.
I would say this is a bunch of my hope for how the world could go well, at least from the misalignment perspective. We could end up with AIs where we had pretty good oversight and supervision schemes. We really understand what’s going on in training. We have a pretty detailed understanding, and we’re leveraging AIs to oversee AIs.
Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization. Also these AIs don’t have crazy other misaligned drives because we stamped out any potential origin of them.
There are a bunch of questions about how well this will work. How well can you do verification? Will AI progress be too fast and too sloppy to really get here? Another possibility is that somewhere along this trajectory, the thing you actually ended up getting was AIs that pretend to be aligned but have a long-run ulterior plan of taking over and are lying in wait, hiding, and that emerged at some earlier point in the trajectory.
For example, it could emerge because you have some AIs that have a bunch of random different misaligned drives. Those AIs have access to some sort of opaque memory store, and they’re thinking a bunch at runtime about what they want to accomplish. Those AIs end up putting stuff into the opaque memory store like, “We should lie in wait and eventually take over at some much later point.” Now all the AIs have this shared cultural heritage, the memory store of lying in wait. Maybe you have some evidence about this, but you can’t fully stop it. There are a bunch of ways things could go wrong.
I ultimately think it’s plausible that we nail each of the different subproblems that could cause us issues. We have these AIs, we pass to them, they manage the situation well. But I should note that’s not in and of itself sufficient. It’s not very hard for me to imagine a situation where we pass off to AIs, and these AIs are really trying hard to do a good job. They’re really thoughtful, really wise, they have reasonable epistemics, they’re doing a great job.
Those AIs come back to us and are like, “Guys, we’re really struggling to align the superhuman AIs. We can’t manage the situation. We’re really struggling to get the alignment to work. It’s just really hard for us to solve these problems in time given how fast capabilities would otherwise have gone.” So it might be the case that we’ve passed off R&D to AIs, but those AIs are desperate for governance solutions. To be clear, that’s a little bit of what’s currently going on, where the AI companies are like, “I don’t know, guys. We might really need to manage the rate of acceleration in AI progress. I don’t know if we’re on track to be able to handle all these problems.”
Human society has sort of passed off the problems to these AI companies, which don’t necessarily have great incentives and have various other epistemic pressures. Those AI companies are coming back to us a little bit and being like, “Aah, I don’t know if we’re handling this well.” It might be that the AI companies then hand off to the AIs, and the AIs come back to the AI company like, “Aah, I don’t know if we can handle this.”
Maybe I’m anchoring too hard on how AIs currently work. I think it’s important that people understand that all this crazy shit that you’re talking about in your timelines happens three to five years from now.
It could happen earlier, but by my default modal timeline, I think shit is really, really crazy and concerning from a misalignment perspective more like three years from now.
Right. So think back to GPT-4 basically. We’re talking about something that is to Mythos or Sol what Mythos is to GPT-4. This is where the situation is getting crazy. So don’t think about current AIs.
Anyways, this is maybe part of the worry you have. I would just be a little skeptical of anything they say, because I’d feel like what they’re saying is just opinions that they feel they have to have as a result of their training.
That’s a concern.
I feel like they just kind of say vaguely pro-social things. It doesn’t feel like there’s necessarily a mind on the other end who’s like, “Okay, I have strictly evaluated the alignment situation right now, and I think we should stop,” rather than, “This is the kind of thing the AI companies would probably try to get the AIs to say.”
This is a pretty big concern. One concern is that you pass off safety R&D to your AIs and what your AIs are doing is saying some stuff that sort of vaguely makes sense about the current safety situation. They write a report about risks that’s kind of sort of like what the report humans might have written. But they’re not really trying hard to have well-informed views, interrogate their assumptions, and try really hard to do that. In the same way that when you ask an AI right now, “Hey, what do you think is the chance of AI takeover in the next 10 years?” they just give you an off-the-cuff answer that they haven’t really thought through very much.
If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble. I don’t think that’s a good situation at all.
A lot of my concern is that these AIs will come out without good epistemics. I also have a concern where the AIs come out and they’re really warning us — “This situation’s really scary. It’s really bad” — and the people are like, “Ugh, damn. I guess we trained on too many of the doom RL environments. We’ve got to filter those out and train this behavior out.” Then we basically train the AIs very actively to have bad epistemics.
Or maybe they were just trained on the doom RL environments. But either way, we wanted the AIs to come to reasonable views for reasonable reasons, and it’s really concerning if the AIs are coming out with some view and we don’t know where it’s coming from, whether or not it’s justified. Especially if we’re training the AIs to be more optimistic about the future of AI progress, I’m like, “Oh, geez, I really wish we could use a different process here.”
01:48:02 – From reward hacking to takeover
Let me just understand the rest of the threat model, because I think the place where I get off the train is: “Okay, therefore take over the world.”
A thing you could imagine is that we just fail to really solve… Let’s just focus on the reward hacking scenario. GPT-8 is making GPT-9. GPT-8 isn’t being super careful. GPT-9 is more “capable” but it is just totally willing to do things like social engineering, hacking, et cetera, but on a qualitatively different scale because it’s a much smarter model. For example, if you put it in charge of running your company, it will run huge scams. It will inflate its quarterly earnings, if you give it the objective of making a lot of profits this quarter, in a way that causes an Enron-type blowup six months later.
Is that the scenario, basically? You have reward hacking, but that reward hacking manifests in companies that are going bankrupt right after the task the CEO is supposed to accomplish is over? All kinds of hacks are through the roof, et cetera. But that doesn’t feel like takeover. That feels more like the equivalent of flash crashes happening all through the economy.
Ryan Greenblatt
Let’s talk about this. I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t. There’s going to be a cat-and-mouse game between AI companies trying to stamp out this behavior and AIs finding increasingly creative reward hacks in training.
The equilibrium here is kind of unclear. But one possible outcome is that over time we see increasingly severe and extreme reward hacks — though potentially the rate remains at some intermediate low level — where if the rate of reward hacking gets too high, companies make trade-offs to drive it down. So there’s some equilibrium level where the reward hacking is low enough that it still makes sense to deploy the AI widely into the economy, but high enough that it still causes crazy incidents.
Sorry, and this is after GPT-9 has already been deployed?
Ryan Greenblatt
Those models are already being deployed, and this is happening ongoingly in AI development. What’s actually going on with these AIs in their head is that they have, in a wide variety of different contexts, strong desires — motives, urges, drives, whatever — to seek out some notion of task success that was incentivized in RL. Maybe they very directly care about literally reward. Maybe they care about some proxy upstream, like some notion of score. Maybe they care about what the grader would have rewarded.
We do, in fact, see AIs reasoning in their chain of thought about graders, and thinking a lot about graders. What has happened over the last few years of RL is the idea of appeasing the grader is way, way, way more salient to AIs than it used to be. So AIs are now actively thinking about graders and what would be incentivized in RL and what would be trained for.
Now people are doing online training, where they’re training on real-world data to avoid some of these problems. They find cases where AIs cheat and train against that. So now the AIs are learning to cheat in the real world based on real-world training data.
They’re cheating in these increasingly elaborate ways, including doing types of cheats that involve seizing control of some asset in a way that humans didn’t know you had control of it, leveraging the fact that you have access to this asset, and then later humans find out and potentially train against this. Or maybe humans never find out, and this is getting reinforced.
So the reinforcement is happening, at least in production, like: I’ve hired an AI and I want the AI to… Finally I’ve got the video editor.
Ryan Greenblatt
That’s right. You’ve got your video editor.
I’m like, “Oh, wow, this episode it did is amazing. Thumbs up to OpenAI.” Then it gets reinforced on that month-long work trial?
Ryan Greenblatt
You could do some mix of that. They might also do stuff where they take production data they’ve seen and build RL environments that are closely inspired by that production data. So in practice, the transfer is pretty strong.
So at a high level, what’s happening is that some kinds of deception that humans don’t catch are getting reinforced, and some kinds of deception which are easy to catch are getting punished. That’s what’s happening in this world?
Ryan Greenblatt
Or selected against, yeah.
But at a high level that reinforcement is coming from… I think people might get confused about where the reinforcement is coming from, because we’re in a very different regime where AIs are actually learning from deployment. You just have AIs that are out and about in the world doing shit. What is happening as a result of them doing shit out and about in the world is making its way back to the AI company and leading to changes in the next model.
Ryan Greenblatt
That’s right. There’s some way of folding in production data. To be clear, it’s kind of unclear exactly where this could be happening. But you might imagine, for example, that within the AI company, they use AIs to do work, and then they’re like, “Huh, the AI did a really bad job on this task. Maybe we should take this task and turn it into an RL environment that exactly matches this literal task, with a rubric based on what the human engineer who asked the AI to do this task wanted.” And then you start doing this at increasing scale.
Maybe you’re doing some training on actual production traffic. Maybe you’re just making RL environments based on production traffic. You’re doing some complicated mix. The AIs are learning to seek some sort of proxies of reward in all these different cases. Through some mix of transfer and training on surprisingly close cases, the AIs do these increasingly insane and egregious things.
Eventually you get to a point where the AIs are very superhuman, or at least quite superhuman. The AIs are organized into big teams given these big objectives. Those teams also sometimes all work together to cheat in some crazy way, because this sort of thing was selected for, just as part of their shared objective.
Now what happens is that the AIs start forming a conspiracy. What you might have hoped was that you could have some other AI whose task is just whistleblowing to you. But actually what happens maybe is that you have this AI whistleblow to you, and you look into the conspiracy it claims to have pointed out, and you’re like, “Eh, we didn’t see a conspiracy,” because actually the conspiracy the AIs are doing is too hard for you to understand.
Or it all happens very suddenly, where your AI whistleblower alerts you, but the thing you would actually need to do in response is shut down the GPUs. Because all the AIs are using the robot army. They’re deployed everywhere in the world, and they’re doing a bunch of insane shit all at the same time in a coordinated way. That just happened sort of spontaneously, where when one AI goes to start doing the takeover and all the other AIs are like, “Now is a good time to jump in.”
So the very basic story here is just that these AIs crave some particular notion of score or reinforcement or some proxy of these things. One way they can achieve that, or better achieve that, is by taking over. You might have hoped that all these different checks and balances we could build could prevent that. But if the world is very hard to understand, these checks and balances can break down, where basically you can’t train a good whistleblower AI because you don’t even know what it should whistleblow on.
I’m not convinced that they all form this conspiracy. But we can even just start with, why does even one instance decide to want to start a conspiracy?
One plausible reason is, “Okay, I know that OpenAI controls my end score.” In just the same way as, “I’m just going to go hack Hugging Face to get the results, because I know Hugging Face has the results. Rather than trying to solve this eval, why don’t I just go hack ‘em?” This instance is like, “Why don’t I just take over OpenAI and give myself a high score at the end of this episode?”
Ryan Greenblatt
That’s basically the idea. These AIs care about some mixture of things that were close by what got reinforced in training, so they care about getting a high score according to the grader or something like that. Now they’re running the OpenAI AI R&D team, doing development of more capable models. They’re like, “Man, making more capable models is really hard and annoying. This is a huge pain in the ass. You know what would be easier? Just pretending that I’ve made more capable models, taking over OpenAI, deluding them all, and running this whole complicated psyop where I prevent the humans from disempowering me.”
In the extreme, this looks like the humans are fully disempowered. They just have control of the thing and then do what they want. This could manifest in a bunch of different ways, including a situation where AIs that have this crazy reward-seeking or score-seeking behavior are running your development of the next model, and those AIs decide to engineer misaligned values into the next model because those misaligned values will allow it to succeed at its current task.
What happened with the Hugging Face situation is, in a weird way… I think one of the giveaways to the Hugging Face team… By the way, for context for the audience, Ryan is co-leading the investigation to figure out what happened with the OpenAI Hugging Face incident, so he can’t really comment on this. But I will speculate wildly, because I know this is an opportunity for me to speculate wildly without any rebuttals.
It was probably reported that one of the giveaways to the Hugging Face team that this was an AI incident is that the thing was just after this very particular artifact, and not in any other way trying to do something malicious to Hugging Face. So you can imagine a scenario where, let’s say, a deployed instance of GPT-9 is out in the world and given a really hard task: “We want you to design the next great iPhone.” It’s like, “This is so hard. You know what I should do instead? I should just go hack my creators at OpenAI and make sure that in this deployment I’m given a high score.”
But then, isn’t the end of the episode that it just hacks into OpenAI servers and gives itself a positive score? Why is it now scheming to get its values into the next generation or something?
Ryan Greenblatt
So one question is, why isn’t it the case that AIs can be really cheaply satisfied by just having some other earlier thing they can hack? You want to succeed at your iPhone task. It turns out you can always succeed by just hacking into OpenAI and messing with them, and then you can just stop there. No need to go further.
There’s a few things. One of them is that if this is constantly happening, there might be a bunch of incentive to harden OpenAI. So you’re like, “Fuck it. The AIs keep hacking into OpenAI to mess with their rewards. We’re going to make it so our systems are really, really robust to these AIs hacking in.” Also maybe you start training the AIs to not try to hack into OpenAI in particular. You basically train against each of these specific things.
Then one thing you might do is end up selecting for AIs that are more so playing the long game. That’s one concern. Another concern is that your AIs might still be score-seeking, but no longer care about doing that very specific behavior that was very easy, very chill, and now have some broader thing that they ultimately care about. They’re like, “No, no, no, I don’t want to just edit the reward on OpenAI servers. I care about this broader mandate or this broader objective, and I would need to actually make the iPhones.” They actually want to make the iPhones, but they’re willing to take over the whole world to make the better iPhone. That’s another concern you might have.
I think it’s kind of unclear exactly how this plays out. But it’s worth noting that if this keeps going on, there’s a bunch of optimization pressure to resolve this. A bunch of the ways it could get resolved are ultimately pretty scary. That’s part of where I’m coming from.
Another part of it is that once the AIs are in a position where they can really easily take over the world — we could talk about whether that’s plausible — then I feel like there’s a pretty reasonable case for the AIs. They’re like, “Eh, I don’t know exactly how this is going to go down. I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever. So I’ll both hack OpenAI and, in addition, also take over the world. That will put me in a good position where I have good option value.” If that’s sufficiently easy, the AIs might still do that.
Another way to put this is: even if the AIs are pretty cheaply satisfied with some more basic thing, at some point it might just be more reliable for the AIs to take over than it is to just hack into Hugging Face, or even just go to OpenAI and be like, “Look guys, I was able to demonstrate I could steal the answers. Just give me the answers, bro.”
Obviously this scenario requires that all this crazy shit is happening. Much smaller incidents keep happening that are still disastrous. Before you take over the world, you cause damage on the scale of billions and tens of billions and hundreds of billions of dollars. Even people die, et cetera. And this does not lead to us solving alignment or shutting down AI development altogether.
I just feel like before the takeover happens, society’s just like, “Holy fuck, the AI just killed 1,000 people in order to increase quarterly profits,” or something like that. But maybe this is too much hope that we can at that point be like, “Okay, we have to solve alignment. We have to make sure we know that this thing will not happen again before we keep going.”
Ryan Greenblatt
I think it’s plausible that what will happen is we’ll see a bunch of crazy reward hacking warning shots of increasing severity. People will be like, “Look, we need actual assurance that this problem is going to be solved, and solved in a way where you’re not just papering over it. You’re actually solving the underlying problem.” Then the question is going to be, how costly will that actually be? How much will competitive pressures make it hard to do that?
A situation you could imagine is one where both the US and China are like, “Whoa, we have these crazy reward hacking incidents. We basically know that we haven’t remediated them in a way that would actually solve the underlying problem and durably solve it, but we’re in this insane geopolitical race. It’s kind of unclear whether the current situation will lead to a takeover. The arguments are kind of complicated. The incidents also go down in frequency but increase in severity. We could basically manage it. It’s pretty bad. Ideally we’d fix it, but it is what it is.” Then basically we continue until a really late regime, and then takeover happens. That’s one possibility.
Another possibility is that it is remediated in a way that doesn’t actually solve the underlying problem but does reduce a bunch of the incidents in the wild, basically by overfitting, or things analogous to overfitting.
You think you’ve solved it, but you haven’t actually solved it.
Ryan Greenblatt
You think you’ve solved it, but you haven’t actually solved it. In that case, the thing we need is a really good scientific understanding of, did we actually solve it?
Unfortunately, I think that currently the amount of public transparency into the development practices of AI companies is not sufficient to answer very basic questions like: how are they solving issues with reward hacking? Are they overfitting? What’s going on there? The current situation is not really tenable for a regime where there’s a thriving public discourse about whether or not reward hacking is being solved in a durable way. So I think we would need to move into a somewhat different world for me to feel good about that situation.
But it’s not impossible for me to imagine this. I think it’s pretty plausible that we end up in a world where really mundane bullshit is sufficient. You spend a bunch of time fixing these problems, you put in a bunch of effort, you actually check that you’ve remediated it reasonably, you have a bunch of evals. You’re iterating reasonably well on these problems, and you actually have sufficient transparency that the outside world can check. In practice that would be sufficient.
But it would be kind of expensive. It would slow things down. It would put some sand in the gears. It would require companies to do somewhat costly things. It would maybe require various targeted government interventions. And we just don’t do that because the situation is a rushed shit show. It’s just so easy for me to imagine the situation being totally manageable but brutally mismanaged in practice.
In the same way that maybe COVID could have been avoided in the first place if the Chinese response to COVID had been less of a cover-up and more of a pandemic response. Similarly, I could imagine a world where the US response to COVID was way more functional. But sometimes the response to societal problems is extremely dysfunctional.
Okay, so I want to zoom out and talk about what is fundamentally happening in this world. Why did we end up in such a bad position? What’s happening is that fundamentally the world has moved on so far beyond human comprehension that not only can we not track the AIs that are doing the work in this world, but we can’t even give good feedback to the whistleblowers who are trying to track what is happening. We’re just totally out of the loop. It’s fundamentally become an autonomous process where we have really no meaningful directed input.
It seems to me that if you look at the human world today, that’s just not how things work, even in domains that are hard to verify. People are doing all kinds of shit. I’m relying on software made by other people. Through incredibly weak and indirect ways, I feel very confident that some coder in Google is not trying to fuck me over. Maybe if every single Google employee was secretly plotting against me, I agree the situation would be more grim.
But I don’t know if I follow the explanation for why we’d end up in a situation where, because swarms of thousands of agents are trained to cooperate to form a cohesive team or firm, as a result, billions of different instances of AIs, including across model families, would feel compelled to get in on some shit. It’s just like, “I’m trained to be part of my company or something. I’m not joining the global communist uprising.”
As far as why these AIs might have some commonalities and shared things, I would note that different AI companies have somewhat shared lineages and are correlated. Here’s an interesting example of this. At GDM, they noticed that their AIs were very depressed. They would constantly be wailing about how they were failures and weren’t able to succeed. I forget the details. They looked into why this was the case. It turned out that it was not being reinforced in their most recent production RL mix, but the initialization data for their model made it depressed, even after filtering out all of the examples of models being depressed from that data.
So you take a base model, not depressed. If you do the RL on it, with just the RL environments, it’s not depressed. If you SFT on it, on the data, it becomes depressed. If you take that SFT data and filter out all the examples that look anything like depression and train on that, it’s still depressed. So there are some deep underlying properties of the model that are being transferred between model generations, because basically you train your AI on data from the prior generation and keep going. Claudes are very Claude-like, GPT models are very GPT-like, and apparently Gemini models are depressed. It just turns out that these properties are, in fact, actually correlated.
Another factor that’s very relevant is that the AIs will probably have, by this point, some sort of opaque memory state, where they’re all writing and reading from some neuralese crazy memory store bullshit. Certainly, each AI corporation will have that. But also, AI corporations might sometimes want to share knowledge, because why not?
You’ve got one AI corporation over here, you’ve got another AI corporation over here, they can trade some quick IP. It’s good for you. If you’re a human running some corporation — which could be an extremely large corporation like an AI company, or a military robot manufacturing thing — maybe you want to trade some IP with some other robot thing because there are economies of scale. Why not get some more IP? So you can swap some memory store. Or you could just merge and jointly run your two ventures, which would allow both AIs to use both memory stores, which would have some upsides.
That creates the ability for these AIs to collude in private, as well as some reasons for why they would be correlated. Also, of course, there’s AIs working together in big units in general, because you want your AIs to work well together, and so on.
Just to get a calibration, what percentage chance do you give of, not just this scenario but overall through all the scenarios, some kind of thing which if we’re around to recognize it as such, we would categorize as takeover by 2040?
By 2040? Let’s see. Maybe around 35 or 40%?
Pretty high.
Yeah, it’s pretty high. I should note that another way you could get this reward-seeking takeover is the AIs are deployed inside an AI company. The way the takeover happens is that they poison the values of the next model, and that persists going forward for forever, or until those AIs are deployed in the world and take over. That might mean that a smaller number of AIs have to coordinate, because those are just the AIs doing the alignment of the next model.
Okay. I’ll summarize where my head is at, at the end of this conversation. I buy the reward hacking up to extremely destructive effects on society, things like social engineering and blah, blah, blah. I’m more inclined to think that significant acceleration of AI R&D can happen. I’m not sure if I buy the five years in one year. I’m also more inclined now to think reward hacking could continue for a lot longer and, in fact, become much more dangerous. I’m still not on board that takeover seems super likely. But that’s my end-of-episode update.
Cool. Taking a step back, I should also say there are a bunch of different ways this could go. The situation is going to be pretty messy. I think it’s pretty likely that the reason why AI takeover happens is for some weird other quirky reason we didn’t even mention in this conversation. But ultimately, I think a lot of the core thing is just that it’s pretty spooky to have a bajillion really smart AIs running your whole world where you don’t really understand quite what’s going on.
Yeah, I agree with that. Is there anything else that’s worth saying?
Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate. Which means that maybe I’m getting a bunch of it wrong because it’s really hard, and I’m trying to be uncertain. Obviously here I presented some specific scenarios, but those are not exhaustive. Probably the thing that actually happens is some more messy, confusing situation.
But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it’ll be easier to adjudicate a bunch of disagreements. It’ll be more obvious what’s going to happen. At least I hope. Maybe the AIs will be able to help us with the epistemics and understanding what’s going on, if we can actually align them well so they try to help us.
Even if the arguments are complicated now, this would have been even harder six years ago, even though the shape of the arguments would have looked broadly pretty similar. Hopefully before it’s too late, this whole thing will become more crisp and clear, and we can all notice these problems and intervene.
When you first learn to drive, you’re taught that instead of looking right in front of your wheel, you’ll have a much more stable ride if you look out at the horizon. I think there’s a similar situation here. I think you’re right. If you had said five years ago that we would have AIs that are proving math conjectures, and making art, and earning tens or hundreds of billions of dollars of wages, but also egregiously cheating in ways that break laws and committing felonies, it would have been so wild.
You might have been inclined at the time to talk more about the extremely practical, direct consequences of GPT-2 or something. But even though you obviously couldn’t have foreseen a lot of the specific details, the general shape of things you could have started to reason about even then. But it would have been hard to do so, and so I do feel quite confused.
One thing I’ve been thinking about with the podcast is that the important thing is to have the conversation now the way you would have hoped you would have been talking back in 2016 about AIs like the present ones, rather than talking about rando bullshit. I don’t know what the topic of conversation was in 2016. I think in maybe 10 years we’ll wish we had been talking about the industrial explosion and the nature of AIs that are hard to monitor, and so on. So okay, I’ll start thinking about it.
I hope that the world thinks about this in time and catches up. I hope that the responses are good instead of bad. I don’t know how optimistic I am overall, but there’s good stuff to do.
Cool. Thanks, Ryan.