当 OpenAI CEO Sam Altman、Anthropic CEO Dario Amodei、Google DeepMind 联合创始人 Demis Hassabis 以及 SpaceX 负责人 Elon Musk 在周末大致同意放缓 AI 发展时,怀疑者立刻察觉到了别有用心的动机。
这些 AI 巨头宣称,他们的目标是“为前沿发展定速”,至少部分支持了一项提案,内容包括引入第三方审计机构、监管国内实验室,以及达成一项全球放缓协议。然而,批评者认为,他们只是想阻止潜在竞争者、削弱开源运动,并规避真正的法律保障——一些人直接将其称为“卡特尔”。
据业内多方消息源称,真相更为复杂。Amodei 在一篇文章中提出的三步提案,呼吁进行 AI 安全倡导者长期主张的变革。尽管它可能成为监管的替代品,但在特朗普治下,实质性监管本来就不太可能。但专家表示,在一个只会愈发重要的问题上,AI 领导者并不是引领这场行动的最佳人选。“
整个行业需要新的领军者,”纽约大学兼职教授、美国国土安全部前新兴技术政策主管 Nick Reese 表示。“我们一直把 Dario Amodei、Sam Altman 和 Elon Musk 这样的人视为最接近问题、每天都在为此工作的人,认为他们最了解情况。
但事实是,对于我们要建设的目标,从来就没有过一个现实的愿景。”对 AI 进展的担忧已持续升级数月,主要导火索是有关大批智能体在领先前沿实验室 Anthropic 和 OpenAI 眼皮底下发动恶意黑客攻击的披露。两家公司发布的报告火上浇油,Anthropic 研究员 Jacob Coxon 的辞职同样如此,他发布了一封公开信解释自己的选择。“
构建 AI 的人真心相信,它可能在这个十年结束前杀死我们所有人,”他写道,并补充说 OpenAI 和 Anthropic 都没有“负责任地行事”,而是“径直冲向可自我改进的超级智能,拿我们的生命赌博”。在业内,关于 AI 的可怕警告并不新鲜,但 Coxon 的这封信似乎突破了圈层。
仅在 X 上,它就被浏览了超过 1.7 亿次,而 Coxon 和他的故事也登上了报纸和电视。 然而,对相当一部分 AI 安全研究人员和 AI 非营利组织工作者来说,这不过是在让一个早已被广泛讨论的问题获得更多关注。“
很多人把这当成 Dario 的想法,或者认为它来自那些 CEO,但这是错的,”前 OpenAI 员工、现领导 AI 研究非营利组织 AI Futures Project 的 Daniel Kokotajlo 说。“公司之外的人多年来一直在呼吁这件事……一直有越来越多人齐声说,‘请不要这么快就构建超级智能。
我们还没准备好。你们需要慢下来。’”Kokotajlo 说,在人们多年“冲他们喊话要求这么做”之后——包括超过 1,000 名 AI 实验室员工签署了 7 月的一封公开信,该信在 OpenAI-Hugging Face 事件后呼吁放缓 AI 开发——这些 CEO 现在“是在屈从于这种压力,同时还把功劳揽到自己身上,[尽管]这并不正当。”
总体而言,少数消息人士告诉 The Verge,他们认为这些 AI 领导者达成的口头协议是朝着正确方向迈出的一步。NYU 的 Reese 说,他“不认为这全是空话”。Apollo Research CEO Marius Hobbhahn 称这是“一个好主意”,并表示“如果它真的发生,那会是长期以来对安全最有利的事情之一”。
The Midas Project 的 Tyler Johnston 说这是一个“好迹象”。Redwood Research CEO Buck Shlegeris 说这是“一些好消息——显然很难知道这是否会转化为任何真实的东西,但我感到谨慎乐观。””
不过,所有人都同意,要让这一口头承诺成为现实,仍有大量工作要做,尤其是在使其成为一项铁板钉钉的协议方面。许多人仍然质疑 AI 领导者的动机。长期以来,规模更大的科技行业一直通过游说制定自己偏好的规则或承诺自我监管来抢在监管之前行动。
大型平台提出的政策可能会对较小的竞争对手造成更沉重的打击,并用利他主义的言辞来为自私自利的目标辩护。它们被指责进行“安全粉饰”,即做出毫无意义的改变,给人以拥有实际安全保障的虚假印象。人们担心这种情况也会在 AI 行业发生,这并不令人意外,尤其是因为 AI 实验室的自愿安全框架多年来一直受到批评。
多位消息人士认为,对“安全粉饰”的担忧是有道理的。“有一个严重的担忧,那就是他们实际上并不会放慢脚步,”Kokotajlo 说,并补充说,令人担心的是“他们只会引入一些外部审计人员,做一堆安全文书工作——其中一些确实是有益的——但归根结底,这实际上根本不会让他们放慢多少。”
纽约大学的 Reese 将这一策略比作十年前社交媒体平台的套路,当时这些公司开始呼吁采取监管行动,以抢在即将到来、对其更不利的法律之前。Reese 说,对于 AI 行业而言,“这把锤子可能不会在本届政府任期内落下,但我认为,如果下一次选举后出现一个民主党政府,那将非常有可能。”
Tech Oversight Project 执行董事 Sacha Haworth 表示,那种监管将至关重要。她说,任何自愿框架本质上都是监管俘获,“我们不应该让狐狸来管理鸡舍。这不是国会再次将责任外包给行业的机会。”
Political Integrity Project 联合创始人 Daniel Lobo-Lewis 表示,自愿监管很可能会重蹈 Meta 那个基本上没有牙齿的 Oversight Board 的覆辙。然而,大多数人同时也认为,除了 AI 实验室在特朗普政府执政期间已同意遵守的模型发布前审查期之外,各公司目前并无迫在眉睫的监管风险。
特朗普总统周一发帖称:“AI 唯一需要的控制或‘护栏’,就是一位强大而聪明(高智商!)的总统,而美国恰恰拥有这样一位,而且绰绰有余!”他还在 Nvidia CEO 黄仁勋出席一场会议登台时给他打了电话——特朗普通过免提对现场观众说,近期对 AI 的担忧是一场“骗局”,并且“机器人不会接管世界”。
至于监管可能拖累其他公司或开源开发者的担忧,多位行业专家对 The Verge 表示,要求放缓的呼声几乎完全针对那些以精确规模指标界定的大型前沿 AI 实验室。在政府 AI 监管迟迟未至的情况下,最好的选择或许是推动各公司达成一项即时、可衡量且可执行的协议。
例如,Kokotajlo 的 AI Futures Project 提议,AI 实验室应允许审计人员访问其算力预算——并承诺大幅削减用于研究的算力预算,从而减缓 AI 的进展,让其他 AI 实验室得以追赶。(Amodei 的文章已经呼吁 AI 实验室允许外部第三方审计机构——如 METR、Apollo 和 Redwood Research——在一定程度上嵌入其组织内部,并有可能标记和举报有问题的发现。)
可以说,AI 放缓面临的最大单一挑战可以用一个词概括:中国。 AI 领导人和政客长期以来一直将中国定位为美国 AI 进展不能放缓的理由——因为无论这项技术可能变得多么危险,他们宁愿它掌握在美国手中而非中国手中,而中国不会踩刹车。
Lobo-Lewis 将这种对中国的恐惧比作冷战时期的导弹差距。一位 X 用户写道:“如果我注定要死于杀人 AI 之手,我希望它是美国的,而不是中国的。” Redwood Research 的 Shlegeris 表示,以“鲁莽”的方式推进 AI 发展,既不符合美国也不符合中国政府的最大利益,并补充说“这并非没有先例的国际协调水平。”
Johnston 说:“(人们)想当然地认为中国不会合作”,并补充道,“这个问题如此严重、如此广泛,协调解决它似乎符合所有人的利益,就像协调核不扩散符合美国和俄罗斯双方的利益一样。”周一,中国外交部发言人郭嘉昆确实对放缓的呼吁进行了反驳,称其为“散布恐慌”。
Midas Project 的 Johnston 和 Redwood Research 的 Shlegeris 都表示,即使没有中国的合作,鼓励美国进行协调仍然至关重要。而对于 Tech Oversight Project 的 Haworth 来说,放缓是一个“将自己定位为 AI 技术如何开发和使用的指南针”的机会。
她认为反对的论调是熟悉套路的一部分。“每当一个行业想要逃避监管时,中国就会被拿出来当替罪羊。”Kokotajlo 则把这种情况比作他看到过的一个卡通画面:人们坐在一辆驶下悬崖的车里。他回忆说,那个对话气泡里写着类似“万岁,我们领先中国了”的话。
最近所有这些恐慌背后的问题是一个被称为递归自我改进的里程碑,而根据该行业当前的轨迹,我们听到的只是警报声的开始。RSI 指的是一个潜在的行业里程碑,届时 AI 模型可以训练、进步并创建自身的新版本——全程无需人类参与。Anthropic 曾表示这一时点最早可能在 2027 年初到来,而 OpenAI 的首席科学家本月早些时候写道,OpenAI 正在投入大量资源以实现这一目标。
AI 行业越来越多的工程师和研究人员担心,RSI 会先于一系列全新且更严重的 AI 隐患到来,其中包括对社会整体造成更深远影响的更严重网络安全事件。Amodei 将 RSI 列为促使他呼吁放缓的主要因素:“大约从今年夏天开始,AI 的推进速度急剧加快”,而 RSI “正开始在整个行业中出现,包括在 Anthropic……如果不加约束,它可能超出我们理解和控制系统能力的速度,因此即便要推进,也必须极其谨慎,”他在最近的文章中写道。
他提议对 RSI 的速率实施“某种‘限速’”,并将其比作对导弹数量的上限。 在 Jacob Coxon 从 Anthropic 的辞职帖中,他警告会出现“能够入侵任何东西、在一夜之间颠覆任何领域、并获取真正权力与资源的超人系统”。
领先 AI 实验室的许多其他研究人员也呼应了他的担忧。OpenAI 研究员 Jasmine Wang 写道,加速冲向 RSI 的危险程度“怎么强调都不为过”。前 Google DeepMind 员工 Vishal Maini 表示,RSI “如今已迫在眉睫,以至于其他任何选项都不合理。”
对 RSI 的恐惧很容易演变为末日式的想象。“AI 开发者相信他们的技术可能导致人类灭绝(或类似糟糕的结果),”Anthropic 研究员 Samuel Marks 在 X 上写道,“这可能在未来几年内发生。
总体而言,员工越资深,就越担忧。”另一位 Anthropic 研究员兼团队负责人 Evan Hubinger 在 X 上写道,“Jacob 在这里说得对——我们确实真心相信 AI 可能杀死全人类!”(他给出的未来十年内发生概率高于 10%。)
前 Google DeepMind 员工 Alex Turner 写道,“许多研究人员相信他们正在构建某种可能杀死地球上所有人的东西。思考如何阻止这件事,字面意义上就是我的日常工作。” OpenAI 研究员 Micah Carroll 写道,Coxon 的看法是一种“跨党派立场”,在“所有前沿 AI 公司”的研究团队中都存在,而且他们都认为“一切照旧的 AI 开发会带来不可接受的灾难性风险”。
“但是,”Carroll 补充道,“我们也不应把灾难性风险凭空臆想成现实——通过有牙齿的安全要求、国际协调,以及一项共识:除非有足够的安全进展让我们集体确信可以这么做,否则就不构建[人工超级智能],这些风险是可以被大幅降低的。”
这种“国际协调”正是 AI 领袖以及广大公众如今所呼吁的。 “我们从未处于这样一种境地:我们真正理解自己在 AI 方面正朝着什么方向前进,”纽约大学的 Reese 说。“我们一直都有这种模糊而未定义、我们并不真正理解、却必须击败中国才能抵达的终局状态。
当我们甚至不理解路径,甚至不理解终点线是什么——甚至不确定是否存在终点线时,真的很难去竞赛。”
When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to “pace the frontier,” signing on at least partially to a proposal for embedding third-party auditors, regulating domestic labs, and reaching a global slowdown agreement. Their critics, however, argued they simply wanted to stop would-be competitors, kneecap the open-source movement, and avoid real legal safeguards — some dubbed it an outright “cartel.” The truth is more complicated, according to sources across the industry.
The three-step proposal, laid out in an essay by Amodei, is calling for changes long espoused by AI safety advocates. While it could become a substitute for regulation, under Trump, substantial regulation is unlikely anyway. But experts say that on an issue that’s only likely to grow in importance, AI leaders aren’t the best people to lead the charge. “The industry as a whole needs new champions,” said Nick Reese, an adjunct professor at New York University and the Department of Homeland Security’s former director of emerging tech policy. “We’ve looked at people like Dario Amodei, Sam Altman, and Elon Musk as these people who are closest to the problem and working on it every day, and they’re the ones who know best.
But the truth is there’s never been a realistic vision for what we’re building toward.” Concerns about AI advancement have been escalating for months, sparked largely by revelations that swarms of agents were behind rogue hacks going on right under the noses of leading frontier labs Anthropic and OpenAI. Reports from both companies fueled the fire, as did the resignation of Anthropic researcher Jacob Coxon, who posted a public letter explaining his choice. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, adding that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.”
Dire warnings about AI aren’t anything new inside the industry, but Coxon’s letter seemed to break containment. It’s been viewed more than 170 million times on X alone, while Coxon and his story have made the rounds in newspapers and on TV. To a significant chunk of AI safety researchers and AI nonprofit workers, however, this is simply drawing more attention to an already widely discussed issue. “A lot of people are thinking of this as Dario’s idea, or it’s coming from the CEOs, but that’s false,” said Daniel Kokotajlo, an ex-OpenAI employee who now leads the AI Futures Project, an AI research nonprofit.
“People outside the companies have been calling for this for years … There’s been this growing chorus of voices saying, ‘Please don’t build superintelligence soon. We are not ready. You need to slow down.’” Kokotajlo said that after years of people “yelling at them to do this” — including more than 1,000 AI lab employees signing a public letter from July, which called for a slowdown in AI development after the OpenAI-Hugging Face incident — the CEOs are now “now bowing to that pressure and also claiming credit for it, [although] not rightfully.” Overall, a handful of sources told The Verge they believe that the verbal agreement made by the AI leaders is a step in the right direction.
NYU’s Reese said he “do[es] not think it’s all hot air.” Apollo Research CEO Marius Hobbhahn called it “a good idea,” saying “it’s one of the best things for safety in a long time if it actually happens.” The Midas Project’s Tyler Johnston said it was a “good sign.” Redwood Research CEO Buck Shlegeris said it was “some great news — obviously it’s hard to know whether this is going to convert into anything real, but I’m feeling cautiously optimistic.” All agreed, though, that there was still a lot of work to do to make this verbal commitment a reality, particularly when it comes to making it an ironclad agreement.
Many still question the motivations of AI’s leaders. The larger tech industry has spent years getting ahead of regulation by lobbying for its own preferred rules or promising self-regulation. Big platforms have proposed policies that could hit smaller competitors harder, using altruistic language to justify self-serving goals. They’ve been accused of safety-washing, or making meaningless changes that give the false impression of actual safeguards. It’s no surprise people are concerned this will happen in the AI industry as well, particularly since AI labs’ voluntary safety frameworks have been criticized for years.
Several sources believe that concerns of safety-washing are valid. “There’s a serious concern that they’re not actually going to slow down,” Kokotajlo says, adding that the fear is that “they’ll just bring in some external auditors, do a bunch of safety paperwork — some of which will be genuinely good — but at the end of the day, it actually won’t slow them down very much at all.” NYU’s Reese compared this gambit to the social media platforms’ playbook a decade ago, when companies began calling for regulatory action to get ahead of impending, less favorable laws.
For the AI industry, Reese said, “the hammer may not come in this administration, but I think if there were a Democratic administration after the next election, there would be a really good possibility.” That kind of regulation would be vital, said Sacha Haworth, executive director of the Tech Oversight Project. She said any voluntary framework is essentially regulatory capture and that “we should not be letting the foxes run the henhouse. This is not an opportunity for Congress to once again outsource responsibility to industry.” Daniel Lobo-Lewis, co-founder of the Political Integrity Project, said that voluntary regulation will likely go the way of Meta’s largely toothless Oversight Board.
Most people also, however, believe companies are in no imminent danger of regulation, besides the model pre-release review periods AI labs have agreed to under the Trump administration. President Trump posted on Monday that “the only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” He also called Nvidia CEO Jensen Huang while Huang was onstage at a conference — Trump told the crowd on speakerphone that recent concerns about AI were a “hoax” and that “the robots will not be taking over.”
As for concerns that regulation could hold back other companies or open source developers, calls for slowdowns have been almost exclusively aimed at big frontier AI labs defined by exact size metrics, many industry experts told The Verge. With no government AI regulation incoming, the best option may be to push companies into an immediate, measurable, and enforceable agreement. Kokotajlo’s AI Futures Project, for instance, proposed AI labs giving auditors access to their compute budget — and pledging to decrease their compute budget for research significantly, which would then slow down AI advancement and allow other AI labs to catch up.
(Amodei’s essay already called for AI labs to allow external third-party auditors — like METR, Apollo, and Redwood Research — to embed within their organizations to some extent and potentially be able to flag, and blow the whistle on, problematic findings.) Arguably the single biggest challenge to an AI slowdown can be summarized in one word: China. AI leaders and politicians alike have long positioned China as the reason why US AI progress can’t slow down — because no matter how dangerous the technology may become, they’d rather it be in US hands than Chinese ones, and China won’t pump the brakes.
Lobo-Lewis compared China fears to the Cold War missile gap. One X user wrote, “If I am going to die at the hands of killer AI, I want it to be American, not Chinese.” Redwood Research’s Shlegeris said it wasn’t in the best interest of neither the US nor the Chinese government to pursue AI development in a “reckless” way, adding that “it would not be unprecedented levels of international coordination.” “[People are] taking for granted that China won’t cooperate,” Johnston said, adding, “This is an issue so serious and so widespread that it seems like it’s in everyone’s interest to coordinate on it, in the same way it was in both the US and Russia’s interest to coordinate on nuclear de-proliferation.” On Monday, Chinese Foreign Ministry spokesperson Guo Jiakun did push back on the calls for a slowdown, calling it “fearmongering.”
The Midas Project’s Johnston and Redwood Research’s Shlegeris both said that even in the absence of Chinese cooperation, it’s still vital to encourage US coordination. And for the Tech Oversight Project’s Haworth, the slowdown presents an opportunity to “position ourselves as the compass for how AI technology can be developed and used.” She sees arguments to the contrary as part of a familiar playbook. “China gets brought up as a bogeyman every time that an industry wants to escape oversight.” Kokotajlo simply compared the situation to a cartoon he’d seen, with people in a car driving off a cliff.
He recalled the speech bubble saying something like, “Hooray, we’re ahead of China.” The issue underlying all the recent panic is a milestone known as recursive self-improvement, and based on the industry’s current trajectory, we’re only hearing the beginning of the alarm bells. RSI refers to a potential industry milestone at which point AI models can train, advance, and create new versions of themselves — all without human involvement. Anthropic has said this point could come as soon as early 2027, and OpenAI’s chief scientist wrote earlier this month that OpenAI is directing a lot of resources towards reaching this goal.
Increasingly, engineers and researchers in the AI industry worry that RSI precedes a whole host of new and more serious AI concerns, including more severe cybersecurity incidents that more deeply impact society at large. Amodei cited RSI as a major factor in his decision to call for a slowdown: “since roughly this summer, AI has been advancing drastically faster” and RSI is “starting to happen across the industry, including at Anthropic … Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all,” he wrote in his recent essay.
He proposed implementing “some kind of ‘speed limit’” on the rate of RSI and compared it to caps on missile numbers. In Jacob Coxon’s resignation post from Anthropic, he warned of “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” Many other researchers at leading AI labs echoed his concerns. It’s “hard to overstate how dangerous” it is to speed towards RSI, wrote Jasmine Wang, an OpenAI researcher. Vishal Maini, an ex-Google DeepMind employee, said RSI is “now so imminent that no other option makes sense.” Fears about RSI can easily turn apocalyptic.
“AI developers believe their technology could cause human extinction (or similarly bad outcomes),” Samuel Marks, an Anthropic researcher, wrote on X “This could happen in the next few years. In general, the more senior the employee, the more concerned they are.” Another Anthropic researcher and team lead, Evan Hubinger, wrote on X that “Jacob is correct here—we really do earnestly believe AI could kill all humans!” (He put the chance at greater than 10 percent over the next decade.) Alex Turner, a former Google DeepMind employee, wrote that “many researchers believe they are building something that could kill everyone on the planet.
It was literally my day job to think about how to stop that.” Micah Carroll, an OpenAI researcher, wrote that Coxon’s belief is a “cross-partisan position” with research teams across “all frontier AI companies,” and that they all believe that “business-as-usual AI development poses unacceptable catastrophic risk.” “But,” Carroll added, “we should also not hyperstition catastrophic risks into existence – they can be greatly reduced via safety requirements with teeth, international coordination, and a consensus to not build [artificial superintelligence] unless there are sufficient safety advances to make us collectively confident to do so.” That kind of “international coordination” is exactly what AI leaders, and the public at large, are now calling for. “We have never been in a situation where we actually understood what we were driving towards with regard to AI,” NYU’s Reese said.
“We’ve always had this amorphous undefined end state that we don’t really understand but we have to beat China to get to. It’s really hard to race when we don’t even understand the path or even understand what the finish line is — or even if there is one.”