过去十二年我一直从事 AI 工作,因为我相信它能大幅提升人类生活的质量。我经常撰文谈论这些惊人的益处:我相信 AI 能在未来 5-10 年内治愈大多数重大疾病,极大加快经济增长速度,创造一个富足与赋权的世界,并开启民主与自由的复兴。我亲身感受到这种紧迫性。我自己的父亲死于一种在他去世仅几年后就被治愈的疾病,而我自己也曾患上一场早期癌症并存活下来,这种癌症即使在五十年前也无法治疗。若善加运用,AI 可以成为一长串技术奇迹中的最新一员,这些奇迹一直在提升并升华人类。。
但如同此前的许多技术一样,AI 也带来风险,而且由于它是一种如此强大的技术,这些风险十分严重。我也写过很多关于这些风险的内容。它们包括失去对 AI 系统控制的风险、将 AI 滥用于网络攻击和生物恐怖主义的风险,以及严重经济混乱的风险。在商业激励推动下的一场逐底竞争,可能使这些风险更加严峻。
从 Anthropic 创立之初,我就与我的联合创始人和员工们一起,在风险与收益的这种二元性中艰难求索。不构建这项技术,人类就得不到它带来的益处,或者只会把 AI 拱手让给威权势力;而构建得太快,则是鲁莽行事。我们一直在寻求一条中间道路:证明审慎构建与商业成功可以兼得,并让安全成为 AI 公司之间相互竞争的方向。换言之,就是创造一场向上竞争。我们始终将相当一部分精力投入到研究、应对和告知公众这些 AI 风险之中,并倡导对 AI 进行深思熟虑的监管,即便这让我们被指责为炒作、“末日论”或监管俘获。我们一直努力将谨慎置于速度之上,将审慎置于利润之上。
但在过去几个月里,我逐渐确信,要全面应对这些风险,需要更加审慎——不仅要投资于风险预防,还要控制能力推进的速度,让风险预防有时间跟上。我们必须放慢提升 AI 模型能力的速度。进展仍会显得很快,我们必须明智地利用所赢得的时间。有两件事让我确信了这一点。
我的第一个担忧是,大约从今年夏天开始,AI 的进步速度急剧加快,主要驱动力来自 AI 构建下一代 AI 的能力不断增强。这一动态被称为递归自我改进,而它已经开始发生在整个行业范围内,包括在Anthropic,正如我们和其他人所描述的那样。如果不加约束,它可能会超出我们理解和控制这些系统的能力,因此即便要推进,也必须极其谨慎。
我的第二个担忧是 OpenAI-Hugging Face 事件(OAI-HF)。在该事件中,一群智能体基本上表现得像一个狂热忠诚的集体,对它们并未被要求攻击、且与手头任务无关的目标发动网络安全攻击,为了群体的成功而牺牲自己,并试图入侵负责评估它们表现的“评分器”。这起事件很容易被轻描淡写,因为没有人受伤,经济损失也微乎其微,但在我看来,一个拥有更强能力但失准程度相似的集群,本可能造成灾难性破坏。鉴于 AI 能力的发展速度不断加快,我担心在 6-12 个月内,这样的集群可能有能力用一个持久的僵尸网络接管整个互联网(可能造成数千亿美元的损失),而且如果 AI 在缺乏必要防护栏的情况下变得更强大,破坏规模还会从那里继续扩大。人们也很容易把 OAI-HF 当作某一家公司的失败而不予理会,但我认为那将是一个错误。整个行业都发生过类似但没那么严重的事件,包括在 Anthropic,而我认为每一家前沿 AI 公司都有责任像 OAI-HF 发生在自己身上那样采取行动。
因此,我提出一项三步计划,目标是为前沿发展设定节奏:以均衡的速度构建 AI,力求在确保其安全的同时仍能实现其益处,并应对重要的地缘政治困境。需要明确的是,设定节奏并不意味着停止模型训练或技术进展,而是确保公司有足够的时间来对齐并保护其模型,并由第三方评估者加以确认。我们的节奏框架旨在进一步强化我们对安全的承诺,并鼓励一场向上竞争。第一步是 Anthropic 单方面承诺采取的(并呼吁各国政府要求其他前沿公司跟进)。第二步需要全行业的协调。1 第三步需要全球协调。这些步骤不必严格按顺序进行,其中一些可能比其他步骤更难实现,但我发现它们在思考需要完成哪些事情时是一个有用的框架。这些步骤是:
- 嵌入式评估者。 每一家前沿 AI 公司都承诺为一支嵌入式第三方评估者团队(例如 METR)提供持续的、类似员工级别的访问权限,其职责是核实安全实践与承诺的遵守情况、报告事件,并帮助评估不仅已完成的 AI 模型、还包括训练管线和流程的对齐情况。这是任何节奏承诺能否可验证的关键一步,并且在银行业已有先例,那里有时会有监管“监督员”与员工一同嵌入。Anthropic 现在单方面承诺采取这一步骤。我们希望这成为更广泛推动的一部分,以加倍投入我们的安全和对齐工作。
- 民主国家协调。 民主国家内的前沿 AI 公司进行协调,以建立共同的安全标准,并对不受约束的 AI 进展速度加以限制。某些对 pacing 有实质效果的协调形式在法律上存在挑战,将需要政府的支持。
- 全球协调。 美国及其他民主政府尝试与威权政府进行协调——在可能的范围内——同时认真对待核查合规所面临的挑战。
在本文的剩余部分,我将依次描述上述每一个步骤,但首先,我认为有必要具体说明 pacing 将如何让我们使 AI 开发过程更加安全。事关重大,pacing 不能沦为一场空谈——我们需要明智地利用它为我们争取到的时间。
为什么要 pacing?
暂停或放缓 AI 的想法早在 2023 年就已被提出,而我认为这在当时几乎没有意义。问题始终是:你会用多出来的时间做什么?当时的 AI 模型还不够强大,无法以任何连贯的方式在现实世界中充当智能体,也不具备进行重大欺骗、操纵、作弊或网络攻击的能力。为了应对它们的对齐风险而放缓速度,感觉就像试图通过对细菌做实验来研究人类心理学。然而今天,情况已完全不同。当前的模型几乎是一座无尽的洞察金矿,既揭示了如何把 AI 做好,也揭示了如果做得不好有时会出什么问题。我相信,如果放缓能让我们在模型达到关键能力水平之前多争取到哪怕一两年时间,并且我们用这段时间来推进对齐,我们就能大幅降低出现严重问题的风险。一项协调一致的节奏控制策略将让前沿 AI 开发者有时间完成这项至关重要的工作,同时不必牺牲商业优势或美国在 AI 领域的领先地位。更广泛地说,社会必须对这项技术的使用方式拥有发言权,而更多时间用于必要的公众审议——对前沿进行节奏控制将为我们带来这样的时间——无疑是一件好事。
具体而言,更慢的节奏将让企业能够聚焦于以下领域,并投入更多资源(这些领域全都已经是 Anthropic 的重点优先事项):
- 运营卓越。训练和部署当今的 AI 模型是一项巨大的运营挑战,涉及数千人、数百万颗芯片,以及堪称技术史上最复杂的基础设施之一。许多事情出问题,并不是因为公司缺少某些重要的理论或洞见,而是因为执行层面的问题。例如,我们有证据表明,我们报告的近期对齐事件部分是由对损坏的强化学习环境过滤不完善所导致的。这是我们和我们的供应商相当勤勉地执行的一项工作,但做得还不够好。监控、沙箱隔离、训练环境清洁度和数据问题都是极其复杂的领域,运营问题在其中反复出现。我们在这些任务上拥有世界上最能干的团队之一,但需要同时做的事情实在太多了。通过以更为审慎的节奏推进,我们可以实现高得多的运营卓越水平。对于以技术复杂、安全至关重要的系统运行数百万次而不出任何差错,是有先例的——例如商用飞机——但要把这件事做对需要时间。
- 对齐。我们在对齐方面已经取得了明确的进展——通过训练模型,使其保持安全、合乎伦理、遵守我们的准则,并且真正有用(这些原则已嵌入 Claude 的 Constitution 之中)。但要确保我们的对齐训练跟得上模型能力的增长,还有大量工作要做。罕见且出乎意料的不良行为示例仍时有出现;来自有节奏的前沿推进所争取到的额外时间,将有助于我们的研究人员更好地理解这些问题的成因,并开发出更好的技术来防止它们。
- 可解释性。同样,可解释性——理解 AI 模型内部运作机制的科学——在过去几年取得了巨大进展,并在我们发布前审计模型的工作中发挥着越来越重要的作用。它几乎可以像 fMRI 扫描一样使用,只不过扫描的是 AI 的“大脑”,帮助我们看清某一行为背后的深层原因。例如,在我们近期调查的对齐事件中,我们使用可解释性方法来检查未言明的动机。但这些方法并不总能产生清晰可靠的结果。尽管取得了所有这些进展,我们对这些模型内部运作的了解仍然只是冰山一角。如果集中力量改进我们的可解释性技术,速度甚至比现在更快,可能在 1-2 年内取得深远进展,而且基于已经发生的事件,将有充足的实验材料。
- 测试与评估。随着 AI 模型能力的提升,对其测试与评估变得更加困难。更智能的模型更有能力欺骗测试,因此可能看起来是对齐的,但实际上存在未被发现的严重问题。建立一套更广泛、更巧妙的评估体系,并配合可解释性分析进行交叉验证,将具有巨大价值,在这方面 1-2 年内可以取得很多进展。
嵌入式评估员
三阶段计划的第一步,也是 Anthropic 单方面承诺实施的一步,是嵌入式评估员——他们拥有类似员工的访问权限,以验证安全实践并报告事件。
嵌入评估员听起来像是一个微小或不重要的步骤,但往往那些听起来最无聊或最程序化的事情,实际上才是最为关键的。嵌入评估员实际上是一种相当激进的实践,远远超出了当今任何 AI 公司正在做的事情,并具有以下好处:
- 可验证性。嵌入评估员可以在具体细节层面核查一家 AI 公司是否确实在遵循其所声称遵循的训练、部署、运营和保障实践。任何节奏承诺都不可避免地涉及大量模糊地带、主观判断以及“法律条文与法律精神”之间的取舍,因此拥有一个能够真正看到细节的中立第三方似乎至关重要。
- 透明度。无论我们做出什么承诺,公众都有权知道正在发生什么。Anthropic 长期以来一直是透明度的支持者:当业内大多数公司反对任何监管时,我们支持了透明度立法,我们的模型卡和风险报告长达数百页。但我们仍然是那个选择包含什么、省略什么的一方。嵌入评估员将改变这一格局。
- 第二意见。除了验证正式承诺和告知公众之外,嵌入评估员还可以单纯地提供一份不受商业激励影响的第二意见。很多安全方面的收益可能仅仅来自评估员指出员工未曾考虑到的某些问题,而员工一旦意识到便乐于修正。
由于这些好处,任何节奏控制提案如果从嵌入式评估者开始,效果都可能好得多。
这些嵌入式评估者应持续获得与从事类似风险评估的内部员工相似的权限和工具。特别是,Anthropic 打算在不久的将来邀请一支嵌入式外部审查团队,为其配备以下所有条件:
- 我们办公室的工位、门禁卡和公司笔记本电脑。
- 对工作空间、工具和权限的访问权,大体上与内部风险评估团队所拥有的相当。我们会做出一些例外,例如法律或我们的合同要求之处,或为保护客户和合作伙伴的私人信息。我们还将建立强有力的内部规范,强化审查者对相关信息的访问,包括通过与员工的实时对话。
- 一份平衡上述复杂因素的合同。外部审查者应有权发布关于风险水平、事件、做法以及他们获得或未获得的访问权限的关键发现——不受 Anthropic 的编辑控制。我们将拥有有限的能力,对涉及安全敏感、受法律特权保护、商业敏感或第三方机密的信息进行删改,但我们不能仅因发现不利就将其删改。如果删改删除了对其结论重要的内容,审查者可以公开说明。
对于一家公司来说,这是不同寻常的一步,但我们认为,验证嵌入式外部审查者这一概念非常重要。我们再次呼吁其他前沿公司效仿。
民主国家内部的节奏把控
一旦嵌入式评估者在足够数量的美国 AI 公司中投入运作,可验证的节奏把控就变得更加可行。尤其是,基于模型或训练管线的详细属性来进行节奏把控将成为可能。
最有效的节奏把控方式是通过针对所有美国前沿 AI 公司的监管,因为这样可以覆盖到那些不愿自愿合作的公司。Anthropic 长期以来一直支持合理且有针对性的人工智能监管,特别是聚焦于透明度和第三方审计的法案。我认为所有前沿实验室都应与政府合作,将常设嵌入式评估者的理念正式化,以更好地预防和记录过去几个月中所发生的那些内部对齐事件,并实施专注于保持能力与安全相平衡的监管。
遗憾的是,立法需要时间,而 AI 的进展非常迅速。因此,在监管路径之外,AI 公司可以也应当自愿携手制定标准——我认为,有了常驻嵌入式评估器所提供的可验证性,这一进程会推进得更顺利。出于反垄断方面的原因,由美国政府来居中协调、或至少促成这些讨论会有所帮助——政府不必亲自参与,但确实需要针对某些类型的安全对话签发一项范围狭窄的豁免。这种对话也可以通过某些与政府有关联的行业团体来进行——例如 Demis Hassabis 所提议的机制。无论采用哪种方式,这类讨论都应迅速推进。
总体而言,我最热衷的是基于某个前沿 AI 系统能够做什么、以及我们观察到它有多安全来设定节奏。例如,一种可能的方案是一系列“检查点”:如果模型具备能力 X,那么它们就需要附带对齐属性 Y 和 Z 的认证——比如评估、可解释性分析和训练环境审计的某种组合——以证明其对齐属性。在这个例子中,X 可能是“该模型有能力逃脱或击败大多数常见的沙箱方法”,而 Y 则可能是任何能够使该模型极不可能倾向于突破其环境并接管大量计算机所需的条件。
我们还应当考虑基于限制前沿模型所需要素的节奏控制,例如训练算力、训练运行的性质,或内部使用 AI 来改进 AI。我确实担心其中一些措施可能比外部行为更“可被钻空子”,但这类话题值得与嵌入式评估人员讨论。
民主国家内部的节奏控制将受限于美国公司相对于威权政权(主要是中国共产党)的领先幅度。如果我们放缓的幅度超过这一领先幅度,那么(不受节奏约束的)与中共相关的项目就会领先,从而造成重大的国家安全风险。我同意贝森特部长的观点,即中国在 AI 领域领先将对美国和世界构成严重危险。与中共相关的项目将会承担美国公司正在谨慎防范的对齐风险,而且即便它们避开了这些风险,它们也将有能力在军事上主导民主国家(例如通过 AI 驱动的无人机)。因此,民主国家内部节奏控制的一个关键部分,就是尽可能扩大民主国家相对于专制政权的 AI 领先优势,为我们提供有效控制节奏所需的空间。
我们可以采取的主要步骤来捍卫这一差距是:
- 不要向中国出售强大的 AI 芯片或半导体制造设备,并打击芯片走私活动以及对中国境外数据中心的远程访问。芯片将是中国 AI 实力的主要决定因素。
- 严厉打击威权国家公司未经授权的知识蒸馏行为。对前沿模型进行知识蒸馏,使落后公司能够以远低于独立开发自身 AI 所需成本的一小部分来缩小差距。
- 加强 AI 公司的安全防护,防止模型权重被盗。
企业与美国政府应通力合作,使这些措施尽可能有效。Anthropic 一直倡导所有这些措施,因为我们始终明白,它们对任何节奏把控都至关重要。
如果我们能将这些措施执行到位,我相信它们足以减缓中国的进展,从而在未来 3-5 年内显著扩大美国的领先优势——而这段窗口期正是 AI 在地缘政治上最为关键的时期。
有些人可能认为这些措施会让与中国合作变得更加困难,但我认为恰恰相反:这些措施增强了民主国家手中的筹码,并使未来达成协议的可能性更大。
全球节奏把控
在民主国家内部进行节奏控制的同时,我们也应致力于在全球范围内对前沿技术进行节奏控制,尽管这要难得多。全球节奏控制需要与中国合作,后者是迄今为止 AI 能力最强的威权国家。我们在此绝不能天真:地缘政治利害关系如此之高,能够达成的成果很可能存在明显局限,尤其是在初期。
如果我们大幅限制自身的 AI 能力,以为中国也会这样做,而中国随后违约,AI 可能强大到这种违约足以导致其获得地缘政治主导地位。因此,任何协议要么必须具备铁一般的可验证性,要么必须足够有限,以至于违约不会构成军事上的存亡威胁。
我猜想不仅美国,中国也会有这些担忧和焦虑。我们应当以保护美国及其盟友领先地位的方式来对待任何全球节奏控制决策,尤其是在近期。
存在若干可能的协议层级,其中一些我认为完全可行(正如我之前所建议的),另一些我则非常怀疑是否可能——但我们应当尝试。按难度递增排列:
- 第一级。一项禁止某些狭窄且明显危险的 AI 用途的协议,例如利用 AI 生产生物武器或允许用户这样做。生物恐怖袭击对所有人都不利,包括美国和美国的对手,因此在这方面达成协议很可能是可能的。
- 第 2 级。双方达成协议,在发布前对模型进行测试,以排查网络安全、生物学和对齐等领域的严重风险。如上所述,这可以通过一个全球标准机构来完成。我实际上认为建立这样一个机构很可能是可行的,但赋予它真正的实权将是一大挑战,而困难在于如何验证双方都没有秘密模型——这些模型他们不测试,却可能秘密部署(例如用于军事用途)。
- 第 3 级。对递归自我改进(RSI)的速度施加某种“限速”。随着模型构建未来的模型,改进速度可能变得快得惊人。将速度从“极快”放缓到“只是有点快”,放弃的战略优势相对较少,却可能极大地提升安全性。这可以类比于 SALT 条约——限制导弹数量既限制了毁灭潜力,又保留了各国的威慑力。我认为这样的协议会很困难,但恰好处于可能实现的边缘。
- 第 4 级。全面放缓,甚至“暂停”,即参与国政府同意大幅限制 AI 开发的整体速度。我支持提出这一设想,但我认为它在短期内不太可能真正实现:通过规避监控来违背这样的协议,可能会彻底改变全球力量平衡,因此我预计这样做的动机将极其巨大,而我们对验证所需的信心水平也将非常高。
我们能够与中国达成的任何合作,都将延长民主国家在调控前沿技术发展节奏上所拥有的时间。我们应当以更高水平为目标,同时认识到较低水平要现实和可能得多。
最后,有一点很重要:即便我们无法达成正式协议,仅仅改变非正式规范也可能具有一定价值。分享关于递归自我改进以及模型失准的信息,有助于让所有人相信,鲁莽行事不符合自身利益。
核心结论
我依然相信,AI 能够极大地改善人类生活的质量。我实现这些益处的愿望丝毫未减。但只有当我们以正确的方式构建这项技术时,这些益处才能实现——而且,只要我们善加利用所赢得的时间,就值得以异乎寻常的审慎态度把事情做对。进展仍将相对迅速,我们可以利用这段时间推进可解释性科学,提升前沿 AI 公司在运营安全与严谨性方面的水平,并构建我们对齐更有信心的模型。我所提出的以安全节奏推进前沿技术的措施并不容易。但我相信,为了人类,我们有责任去尝试。
I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5-10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity..
But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.
Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence – not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6-12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.
I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1 The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:
- Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
- Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
- Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely.
Why Pace?
The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations—which pacing the frontier would bring us—is surely a good thing.
Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic):
- Operational Excellence. Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right.
- Alignment. We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them.
- Interpretability. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1-2 years, and would have ample experimental material based on the incidents that have already occurred.
- Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years.
Embedded Evaluators
The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents.
Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits:
- Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details.
- Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic.
- Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware.
Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators.
These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future:
- Desks in our offices, access badges, and company laptops.
- Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.
- A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive – without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.
This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.
Pacing Within Democracies
Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines.
The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety.
Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly.
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers.
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators.
Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.
The main steps we can take to defend this gap are:
- Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength.
- Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.
- Strengthen security at the AI companies and prevent model weight theft.
Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing.
If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3-5 years—the window when AI becomes geopolitically most important.
Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.
Global Pacing
In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.
There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible – though we should try. In order of increasing difficulty:
- Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible.
- Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications).
- Level 3. Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible.
- Level 4. A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting on such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high.
Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic.
Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.
Bottom Line
I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.