WSJ just reported: Anthropic researcher Jacob Coxon is quitting AI because he thinks self-improving models could become uncontrollable by 2027.
“We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already,” he said.
His fear centers on recursive self-improvement, where AI takes over enough AI research to speed development of increasingly capable successors.
He moved from OpenAI to Anthropic specifically for its safety work, yet says even sincere safeguards cannot overcome competition without coordinated restraint.
The most interesting point is, Coxon is leaving despite believing Anthropic takes safety seriously.
Anthropic's $2 tn IPO makes that tension so obvious. Anthropic is simultaneously asking the world to believe 2 things: that increasingly powerful AI can create enormous economic value, potentially supporting a $2 trillion public valuation, and that development may eventually need to be slowed when capability crosses dangerous thresholds.
Those positions create a difficult governance test: will safety commitments still bind when obeying them becomes commercially expensive.