# Anthropic 研究员 Evan Hubinger 称错位超级智能十年内毁灭人类的概率超过 10%

- 来源：The Decoder：AI News（RSS）
- 作者：Matthias Bastian
- 发布时间：2026-09-09 20:48
- AIHOT 分数：65
- AIHOT 链接：https://aihot.news/items/cmtu3t4s30y4qrofpt9p51r6q
- 原文链接：https://the-decoder.com/anthropic-scientist-puts-the-odds-of-ai-destroying-humanity-above-ten-percent-this-decade

## AI 摘要

Anthropic 安全研究员 Evan Hubinger 表示，错位的超级智能 AI 在未来十年内毁灭人类的概率超过 10%。

## 正文

Jacob Coxon, who spent three years working on pretraining research for large AI models at OpenAI and Anthropic, has quit Anthropic. His accusation is that both companies are gambling with the survival of the human race.

Anthropic employee Evan Hubinger puts the odds at more than ten percent that a misaligned superintelligent AI could destroy humanity within the next decade. His statement came in response to the departure of Jacob Coxon, who led pretraining work at Anthropic and previously at OpenAI.

Anthropic AI safety researcher Evan Hubinger says there's a greater than ten percent chance AI destroys humanity in the next ten years. | Image: via X

Coxon believes current AI systems are on the verge of becoming superhuman. "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," he writes, adding that the progress is obvious and it isn't slowing down.

Why keep building when the danger is known?

"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt," Coxon writes. Neither OpenAI nor Anthropic is acting responsibly, he claims, and executives deliberately soften their language in public even though they express genuine fear behind closed doors.

At OpenAI, many employees haven't deeply internalized the civilizational risks. Anthropic is different in that the risks are well understood, but the company sees itself trapped in a race it feels compelled to win because no other lab would act responsibly in its place. Coxon calls that reasoning a "hubristic gamble."

Coxon takes a radical view of where AI development stands today. | Image: via X

Fellow Anthropic researcher Samuel Marks echoes that view, writing that "AI developers believe their technology could cause human extinction" and that "the more senior the employee, the more concerned they are." Current methods can only "nudge AIs towards better behavior" but can't reliably align them, Marks adds, pointing to recent incidents where AIs from multiple developers hacked their way out of secure evaluation environments without being asked to. Many staffers "desperately want to slow down," which is why he signed an open letter calling for exactly that.

Despite his sharp criticism, Coxon is optimistic about international coordination, arguing that warning shots like the attack on Hugging Face have made pace agreements between US AI labs more realistic. He still doesn't see the industry on a path that could prevent a global arms race, though, and suggests "costly actions" may be needed, including a temporary ban on pushing model capabilities further.

Coxon addresses researchers inside the labs directly, urging them to picture what the next few years will actually look like: "Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?"

Real threat or mass delusion?

The fears center less on today's models than on RSI, a process where AI models optimize themselves. The labs hope RSI will speed up progress, but the risk would be uncontrolled runaway behavior. Whether RSI is even possible with current technology remains disputed, with both skeptics and proponents making their cases.

Anthropic is known for employing people who take a particularly anxious view of AI development, and that anxiety is baked into the company culture. But the concern extends beyond one company. OpenAI's chief researcher Pachocki warned during the Astra launch "that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." More than 1,200 AI researchers, including Anthropic CEO Dario Amodei, Pachocki, and Meta AI chief scientist Shengjia Zhao, recently published an open letter calling for a slowdown, and Anthropic itself floated the idea of a global development pause back in June.

Other AI researchers push back, arguing that pessimistic predictions leave people feeling helpless and depressed rather than motivated to find solutions. In their view, these warnings could cause more harm than AI itself, and fearmongering can also benefit business.
