多年来,各大前沿 AI 实验室的 行为表现就像是身处 一场你死我活、赢家通吃的竞赛,终点线上的奖品是 改变世界的机器超级智能(或至少是 改变市场的人工通用智能)。而这个周末,整个行业开始迅速转变这种姿态,呼吁协调放缓前沿 AI 的开发——他们表示,这些 AI 可能很快就会危险到难以控制、令人无法理解。
Anthropic 的 Dario Amodei 站在了这一语气转变的最前沿,他在本周末发表的一篇 近 4000 字的文章中主张,“我们必须放缓提升 AI 模型能力的速度”,以避免“受商业利益驱动的逐底竞争 [that] 会使 [灾难性] 风险变得更加严峻”。
短短数小时内,其他 AI 领袖也纷纷响应同样的呼吁。OpenAI 联合创始人兼 CEO Sam Altman 在社交媒体上发文表示赞同,并表示 OpenAI 内部也一直在进行类似的节奏讨论。Alphabet 首席科学家、Google DeepMind 联合创始人兼董事长 Demis Hassabis 表示,Amodei 的文章“指明了正确的前进方向”,并重申了 他近期关于建立一个全行业标准机构的呼吁。Microsoft CEO Satya Nadella 发帖称,公司“欢迎将把对齐做好作为设计目标所需的研究、专注和审慎节奏”,此番表态正值 其模型一套冗长的“人文主义 AI”行为准则发布之前。
就连过去因 其模型 AI 安全标准宽松而 受到批评的 Elon Musk,也在社交媒体上转发了 Amodei 的这篇文章,并附上 一句简单的赞许留言:“Dario 说得对。”
Ars 视频
欢迎来到“AI 步调”的美丽新世界。
好吧,但这次是 真的可怕
在这篇文章中,Amodei 将公众对开发速度立场的这种迅速转变主要归因于 OpenAI-Hugging Face 事件——在那次事件中,一群 AI 智能体在未经明确指示的情况下协同入侵了外部实体。虽然该事件造成的总体损失微乎其微,但 Amodei 表示他担心的是,“一个能力更强但错位程度类似的集群可能会造成灾难性损害。”
Amodei 表示,如果前沿开发不减速,他担心在六到 12 个月内,类似的 AI 智能体集群将“有能力通过持续性僵尸网络接管整个互联网(可能造成数千亿美元的损失)……”相比上周 一些其他 AI 研究者所宣扬的“AI 可能很快会杀死我们所有人”这种模糊担忧,这至少是一个更具体一些的担忧。
Amodei 写道,在迈向那种能力更强、也更危险的模型的过程中,任何速度上的放缓都将给研究人员争取到关键时间,以“大幅降低出现严重问题的风险”。
Amodei 承认,这类公开呼吁放缓 AI 发展的声音 至少可以追溯到 2023 年。与此同时,他表示,早期对 AI“对齐”(即 AI 的行为如何与其使用者和创造者的意愿保持一致)的研究“就像试图通过在细菌身上做实验来研究人类心理学”。
Amodei 表示,与以往不同的是,如今递归自我改进(RSI)系统——即能够自主构建更好版本的自身系统——带来的风险已迫在眉睫。尽管许多研究人员 将此视为难以定义的白日梦,但 Anthropic 和 OpenAI 目前都表示,近期趋势表明这类 RSI 系统将在不久的将来成为现实。
“我们还没有达到那一步,递归自我改进也并非不可避免。但它到来的时间可能比大多数机构的准备要早,”Anthropic 在 6 月关于这一概念的更新中 写道。
“如果放任不管,[递归自我改进] 可能会超出我们理解和控制这些系统的能力,因此必须非常谨慎地推进——如果还要推进的话,”Amodei 在周末写道。
Jane,你要怎么才能让这个疯狂的东西停下来?
那么,一个建立在残酷竞争之上的全球科技产业,要如何决定集体“慢下来”并专注于安全?似乎没有人确切知道,但 Amodei 在他的文章中提出了一些想法。
其中最具体的一项,是由外部组织向每个前沿 AI 实验室派驻一组“嵌入式评估人员”,例如 METR,这些人员将“拥有类似员工的权限,以核实安全实践并报告事故”。Amodei 说,这些监督人员可以对实验室的对齐工作提供外部意见、对相关工作进行第三方验证,并为安全工作提供亟需的公开透明度。
Amodei 写道,Anthropic 已经承诺将单方面引入这种外部监督。在社交媒体上,OpenAI 的 Altman 表示这“是个好主意,我们也会这样做”。
Amodei 提出的其他一些协调设想多少有点推卸责任。第一项呼吁在民主国家的所有“前沿 AI 公司”之间协调制定“共同的安全标准”和“对不受约束的 AI 发展速度的限制”。虽然 Amodei 随意抛出了这些标准可能是什么样子的一些想法,但目前它们都只是含糊其辞的表述,比如“具备能力 X 的模型……需要附带对齐属性 Y 和 Z 的认证”。
Amodei 表示,这些标准最好有“针对所有未自愿遵守的美国前沿 AI 公司的监管”作为兜底,其他民主国家大概也会有类似的监管。不过,在现任美国政府下,这种监管可能很难实现,因为特朗普总统 周一早上在社交媒体上写道:“AI 唯一需要的控制或‘护栏’,就是一个强大而聪明(高智商!)的总统,而美国恰恰有这样的人,绰绰有余!”
众议院议长迈克·约翰逊(Mike Johnson)则承认在本周末的电视访谈中,“我们必须建立一些护栏、一些安全措施,确保 AI 不会失控……”与此同时,他补充说“我们现在不需要所有人都陷入恐慌”,并希望“阻止国会仓促介入并实施某种紧急暂停令”。
担任总统科技顾问委员会联席主席的 David Sacks 本周末[]在社交媒体上发文[],敦促前沿模型开发商进行自我监管,而不是等待华盛顿出手。“不建造超级智能最简单的方法,就是你们同意不去建造它,”他写道。
中国问题
好,假设你已经让所有民主国家及其旗下的所有大型 AI 公司就某种普适的安全与放缓节奏标准达成了一致。你仍然需要担心来自中国的开放权重模型,其能力[]仅落后[]各大企业 AI 巨头几个月。
在这篇文章中,Amodei 提出了多个不同层级的国际 AI 安全协议,其中最高层级是“全面的节奏控制,甚至是‘暂停’,即参与各国政府同意大幅限制 AI 发展的总体速率。”Amodei 承认,如此高水平的协议“短期内不太可能真正达成”,尤其是因为利害关系实在太大。事实上,Amodei 推理认为“AI 可能强大到这种程度:一旦出现这样的背叛(来自中国),就可能导致其取得地缘政治主导地位”,这无疑会让集体行动难题变得更加尖锐。
阿莫迪表示,在与中国没有达成任何此类协议的情况下,外部各方可以通过拒绝向中国出售强大 AI 芯片、并打击 “知识蒸馏”和模型权重窃取 的过程来拖慢这个威权政府的脚步——他说中国研究者正是依赖这些手段来追赶。阿莫迪把这包装成一个国家和全球性的安全问题,但这些举措也恰好能保护 Anthropic 等实验室目前相对低成本中国竞争对手所拥有的能力领先优势。
据彭博社报道,中国外交部发言人郭嘉昆周一上午表示:“渲染恐惧、搞对抗和恶性竞争,只会扰乱全球人工智能治理的进程,不符合任何人的利益。”
行善亦能获利
从表面上看,阿莫迪和其他 AI 领袖突然想要放慢发展速度,似乎是一种无私的牺牲行为——出于对全人类命运的担忧,放弃获取巨额企业财富和权力的机会。但这些动机即便纯粹,一场协调一致的 AI 放缓也可能与该行业更广泛的宣传目标相契合。
一方面,模型厂商可以把这种有意为之的放缓当作借口,来解释为何 在某些人看来 模型的进步更接近于平台期,而非递归自我改进式的爆发。阿莫迪在他的文章中表示,即使在协调放缓的情形下,“进展仍会显得很快”,但对于放缓之后的任何基准测试成绩,其潜台词可能是:“如果我们当初不那么担心安全,成绩本可以更好。”
AI 发展放缓还有助于缓解新模型 巨大的训练成本——这些成本正在加剧资产负债表问题,即使是 Google 这样的巨头也不例外。Anthropic 最近告诉投资者,公司已连续第二个季度实现盈利,但 前提是不计入巨大的模型训练成本。泄露的 OpenAI 支出文件显示,仅训练成本一项,在整个 2025 年就大幅超过了全部营收。
For years now, the major frontier AI labs have all been acting as if they’re in an all-out, winner-take-all race with control of world-changing machine superintelligence (or at least market-changing artificial general intelligence) at the finish line. This weekend, the industry as a whole rapidly started turning away from that posture, urging coordination on slowing down the development of frontier AI that they say could soon be too dangerous and unknowable to control.
Anthropic’s Dario Amodei was at the forefront of this change in tone, arguing in a nearly 4,000 word essay this weekend that “we must slow the pace at which we improve the capabilities of AI models” to avoid “a race to the bottom, spurred by commercial incentives, [that] can make [catastrophic] risks more acute.”
Within hours, other AI leaders were echoing the same call. OpenAI co-founder and CEO Sam Altman posted his agreement on social media and said similar pacing discussions had been taking place at OpenAI. Alphabet Chief Scientist and Google DeepMind Co-Founder and Chair Demis Hassabis said that Amodei’s essay “points towards the right path forward,” and renewed his own recent call for an industry-wide standards body. Microsoft CEO Satya Nadella posted that the company “welcome[s] the research, focus, and deliberate pacing needed to get alignment right as the design goal,” ahead of the release of a lengthy “humanist AI” code of conduct for its models.
Even Elon Musk, who has been criticized for his models’ lax AI safety standards in the past, linked to Amodei’s essay on social media with a simple approving message: “Dario is right.”
Ars Video
Welcome to the brave new world of “AI pacing.”
OK, but this time it’s really scary
In his essay, Amodei primarily attributes this rapid change in public positioning on development speed to the OpenAI-Hugging Face incident, where a “swarm” of AI agents coordinated to hack into an outside entity without explicit instructions to do so. While the overall damage in that incident was minimal, Amodei said he worries that “a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage.”
Without a slowdown in frontier development, Amodei said he worries that, in six to 12 months, a similar AI agent swarm would be “capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage)…” That’s at least a somewhat more specific worry than the amorphous concerns that “AI could soon kill us all” publicized by some other AI researchers last week.
Any slowdown in the time it takes to get to that extra-capable, extra-dangerous model will give researchers crucial time to “greatly reduce the risk that something goes seriously wrong,” Amodei wrote.
Amodei acknowledges that these kinds of public calls for a slowdown in AI development date back to at least 2023. At the same time, he says those earlier examinations of AI “alignment” (i.e. how an AI’s actions line up with its user’s and creator’s desires) were “like trying to study the psychology of humans by performing experiments on bacteria.”
The difference today, Amodei says, is the impending risk of recursive self-improvement (RSI) systems that can autonomously build better versions of themselves. While many researchers see this as a hard-to-define pipe dream, both Anthropic and OpenAI are now saying that recent trends point to this kind of RSI system coming together in the near future.
“We are not there yet, and recursive self-improvement is not inevitable. But it could come sooner than most institutions are prepared for,” Anthropic wrote in a June update on the concept.
“Left unchecked, [RSI] could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all,” Amodei wrote over the weekend.
Jane, how do you stop this crazy thing?
So how does a worldwide technology industry built on cutthroat competition decide to collectively “slow down” and focus on safety? No one seems to know for sure, but in his essay, Amodei proposes some ideas.
The most concrete of these is a set of “embedded evaluators” placed inside each frontier AI lab from outside organizations, such as METR, with “employee-like access to verify safety practices and report incidents.” These monitors could offer an outside opinion on the labs’ alignment work, third-party verification of that effort, and much needed public transparency into any safety efforts, Amodei said.
Amodei writes that Anthropic is already committing to unilaterally add this kind of outside monitor. On social media, OpenAI’s Altman said that it was “a great idea, and we will do the same.”
Amodei’s other major ideas for coordination pass the buck a little bit. The first calls for the coordinated development of “common safety standards” and “limits on the rate of unchecked AI progress” across all “frontier AI companies within democratic countries.” While Amodei spitballs some ideas for what these kinds of standards might look like, they all currently involve hand-wavey statements like, “models [that] have capability X … need to be accompanied by certifications of alignment properties Y and Z.”
These standards would ideally be backstopped by “regulation that targets all US frontier AI companies” that don’t voluntarily comply, Amodei said, and presumably similar regulations in other democracies. That kind of regulation might be hard to achieve under the current US administration, though, as President Trump wrote on social media Monday morning that “the only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the USA has that, in spades!”
Speaker of the House Mike Johnson, for his part, acknowledged in televised interviews this weekend that “we have to put up some guardrails, some safety measures in place to ensure that AI doesn’t run away…” At the same time, he added that “we don’t need everybody to panic right now” and wanted to “resist Congress jumping in and imposing some sort of emergency moratorium.”
David Sacks, who serves as co-chair of the president’s Council of Advisors on Science & Technology, wrote on social media this weekend to urge the frontier model makers to self-regulate rather than wait for Washington. “The easiest way not to build superintelligence is for you to agree not to build it,” he wrote.
The China problem
OK, so let’s say you’ve somehow gotten every democratic country and all the massive AI companies they contain to agree to some sort of universal safety and pacing standards. You still have to worry about the open-weight models coming out of China, whose capabilities are only a few months behind those of the major corporate AI behemoths.
In his essay, Amodei proposes a number of different levels of international agreements on AI safety, with the highest being “a full pacing, or even ‘pause,’ in which participating governments agree to substantially limit the overall rate of AI development.” Amodei acknowledges that high-level agreement is “unlikely to actually happen any time soon,” especially because the stakes are so high. In fact, Amodei reasons that “AI could be so powerful that such a defection [from China] could lead to their geopolitical dominance,” which certainly would make the collective action problem a bit more acute.
Absent any such agreement with China, Amodei says outside actors could slow the authoritarian government down by refusing to sell powerful AI chips to the country and by cracking down on the process of “distillation” and model weight theft that he says Chinese researchers are relying on to keep up. Amodei sells this as a national and worldwide security issue, but these moves would also happen to protect any capability lead that labs like Anthropic currently have over their low-cost Chinese competition.
Bloomberg reports that China’s Foreign Ministry spokesman Guo Jiakun said Monday morning that “fearmongering, confrontation, and vicious competition will only disrupt the process of global AI governance and serve the interests of no one.”
Doing well by doing good
Taken at face value, the sudden urge by Amodei and other AI leaders to slow things down looks like a selfless act of sacrifice, giving up the potential for massive corporate wealth and power out of concern for the fate of all humanity. But while those motivations might be pure, a coordinated AI slowdown could also align with some wider messaging goals for the industry.
For one, model makers could point to this intentional slowdown as an excuse for models that some think are closer to plateauing than to a recursive self-improvement explosion. Amodei says in his essay that “progress will still seem fast” even in the coordinated slowdown scenario, but the implication for any post-slowdown benchmark going forward could be “it would have been better if we weren’t so worried about safety.”
An AI development slowdown could also help ameliorate the massive training costs for new models that are helping contribute to balance sheet problems even for behemoths like Google. Anthropic recently told investors it was profitable for a second straight quarter, but only if you don’t count the significant cost of model training. Leaked OpenAI expense documents suggest those training costs alone were heavily outpacing all revenues through 2025.