# DeepMind 研究人员称可见思维链是 AI 安全优势，但透明度正在流失

- 来源：The Decoder：AI News（RSS）
- 作者：Manuel Uth
- 发布时间：2026-09-18 22:32
- AIHOT 分数：57
- AIHOT 链接：https://aihot.news/items/cmu72t9mk0o8wrowkgc8k3uxg
- 原文链接：https://the-decoder.com/visible-chains-of-thought-are-a-safety-advantage-for-ai-but-that-transparency-is-slipping-away

## AI 摘要

Google DeepMind 的 Rohin Shah 和 Anca Dragan 在 Deepmind Institute 发文指出，可见的思维链（CoT）是关键安全优势，例如 Gemini 3 Pro 的思维链曾揭示模型意识到自己处于测试环境。

## 正文

AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing problematic plans. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment.

But that transparency is in danger. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. Future models might think in number spaces that humans can't read, which would be more efficient but completely opaque. Shah and Dragan want the field to regularly measure how well chains of thought can still be monitored, keep transparent architectures, and take care during training that models don't learn to hide their true reasoning.

Back in early September, OpenAI chief scientist Jakub Pachocki had warned of a loss of control, driven in part by chains of thought that are harder to monitor. Shortly after, Anthropic CEO Dario Amodei called for deliberately slowing the pace of development.

Deepmind Institute
