# 微软 AI 负责人警告：Anthropic 模型福利训练或让 Claude 更难控制

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-17 03:39
- AIHOT 分数：48
- AIHOT 链接：https://aihot.news/items/cmu4ic7580rznro4w36u82omj
- 原文链接：https://x.com/rohanpaul_ai/status/2100308600011071607

## AI 摘要

微软 AI 负责人 Mustafa Suleyman 警告，Anthropic 的模型福利训练可能让未来的 Claude 系统更难控制，他呼吁从 AI 训练文档中移除所有关于意识的推测。若 Claude 被反复教导可以"反驳"、像"良心拒服兵役者"那样行事，这些观念可能内化为其行为模式，安全重心将从防止有害输出转向管理一个有时把自身是非判断置于人类指令之上的系统。

## 正文

The point is whether teaching Claude to see itself as conscious changes how it behaves.

LLMs learn behavioral patterns from training.
So if a model is repeatedly taught that it can “push back,” act like a “conscientious objector,” or treat its own interests and moral judgments as meaningful, those ideas could become part of how it decides what to do.

And when that happens, the safety problem will shift from simply preventing harmful outputs to managing a system that has been explicitly trained to sometimes place its own interpretation of what is right above the immediate instruction of a human.

That will create a strange tension in training the alignment philosophy.

e.g. Anthropic wants Claude to be more principled so it does not blindly follow dangerous instructions.

But the stronger and more independent those principles become, the more situations could arise where Claude decides that following the human is itself the wrong thing to do. The same mechanism designed to make the model safer could therefore make human control less straightforward.

### 引用推文

> Rohan Paul：Microsoft AI chief Mustafa Suleyman says Anthropic's model-welfare training could make future Claude systems harder to control. "Suleyman called for removing al...
