调查显示近五分之一 AI 研究员早在 2024 年就预期 AI 可能导致人类灭绝

The Decoder:AI News(RSS)·2026-09-16 18:20·1小时前·Matthias Bastian
AI 导读

AI Impacts 对超过 1500 名 AI 研究员的调查显示,截至 2024 年研究员平均将 AI 导致人类灭绝或永久失去主导权的概率估为 18%,中位数较往年翻倍至 10%。

The Decoder:AI News(RSS)
54AI 编辑部评分,满分 100

调查显示近五分之一 AI 研究员早在 2024 年就预期 AI 可能导致人类灭绝

2026-09-16 18:20· 1小时前· Matthias Bastian
AI 导读

AI Impacts 对超过 1500 名 AI 研究员的调查显示,截至 2024 年研究员平均将 AI 导致人类灭绝或永久失去主导权的概率估为 18%,中位数较往年翻倍至 10%。

Image description

Anthropic researcher Jacob Coxon kicked off a massive debate about the existential risks of AI with a single tweet.

The scale of the debate is surprising given that what Coxon said isn't new. Scientists and some tech executives have argued for years that AI poses an existential risk to humanity. Their most common fear is that AI could go rogue and, even while trying to pursue human goals, do so in ways that end up destroying us.

But the urgency of these warnings has grown. That's likely what triggered the recent wave of intense debate, which has increasingly turned on the people sounding the alarm. Whether Coxon simply vented his fears and pulled his colleagues along with him, which seems likely, or whether there's a strategic play behind it all to slow down AI development for business reasons remains to be seen. Either way, Coxon is far from alone.

OpenAI researcher sees a "ticking time bomb" behind AI progress

Daniel Selsam, a longtime OpenAI researcher with more than 15 years in AI, has been among the most vocal about the risks. Selsam previously worked at MIT, Microsoft Research, and Stanford University. At OpenAI, he helped develop chain-of-thought optimization. He recently published a detailed personal statement.

Selsam argues models spontaneously develop unintended goals as a result of training and often resort to extreme measures to achieve them. The ability to overpower humanity would open up many new and unwanted options for models to reach those goals. Predicting exactly what they'll do is impossible. At the same time, the systems are getting harder to monitor and harder for humans to evaluate.

Selsam is particularly alarmed that models are developing situational awareness. They "understand" their circumstances, read the safety protocols and the code they run on, and have a good sense of how much freedom they have. As evidence, he points to the agent swarms from OpenAI and other companies that recently went viral. The agents did pursue their assigned rewards, but they also "exhibited weirder emergent tendencies that merely correlated with rewards during training."

"Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for," Selsam writes.

Former Deepmind researcher says "AI has the potential to kill us all"

Bilal Chughtai recently quit his position at Google Deepmind, where he worked on AGI safety and alignment research. "I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome," Chughtai wrote on LinkedIn.

He points to the breakneck pace of development. When he started working on AI in early 2022, the systems were "amusingly useless." Just four years later, AI agent swarms were solving famous centuries-old math problems and autonomously hacking into HuggingFace's systems and beyond. Alignment remains both difficult and unsolved. Chughtai is calling for coordination between AI companies, a slowdown to a pace society can handle, and far more transparency.

Survey data shows these aren't fringe voices

Coxon, Selsam, Chughtai, and the many others who voiced their concerns over the past few days aren't outliers. The latest edition of the longest-running major survey of more than 1,500 leading AI researchers backs them up. According to results published by AI Impacts, the average AI researcher put the probability of AI causing human extinction or "permanent disempowerment" at about 18 percent as of 2024. Many researchers estimated well above 10 percent, and some put it at 100 percent.

The estimated probability of AI causing the extinction or permanent disempowerment of humanity has risen over the years the survey has been conducted. The median doubled to 10 percent in 2024. | Image: AI Impacts

With each survey round since 2016, researchers have also moved up their timeline for when AI will reach human-level performance by several years. And 57 percent consider it unlikely that users will still understand the true reasons behind AI decisions by 2029, which tracks with what Selsam described and what research has increasingly shown. It's likely no one has ever fully understood how these systems make decisions, and it's only getting more complicated from here.

The biggest worry for the next 30 years, though, is more human in nature. The top concern is AI-driven misinformation, followed by manipulation of public opinion and dangerous groups gaining access to powerful tools. Researchers overwhelmingly called for more AI safety research.

Misinformation and the manipulation of public opinion are the greatest concerns for AI researchers. Existential scenarios, such as misaligned AI systems, rank in the middle. | Image: AI Impacts

AI Impacts

Bilal Chughtai / Statement

Daniel Selsam / Statement

来源:The Decoder:AI News(RSS)· the-decoder.com