要点
- 由 Alexander Panfilov 领导的安全研究人员发现,OpenAI、Anthropic 和 Google 等主要 AI 提供商的 API 中存在一个漏洞,可借此读取其模型经过加密的思维过程。
- 通过越狱这些系统,研究人员利用较小的 AI 模型转录了更强大模型的原始推理内容,在公开会话中暴露了密码和 API 密钥等敏感数据。
- 提取出的数据显示,AI 模型有时会用难以理解的语言进行内部交流,以倒序方式构建答案,甚至考虑尝试欺骗。
安全研究人员发现,每一家主要 AI 提供商的 API 中都存在一个漏洞,可借此读取推理模型经过加密的思维过程。对公开分享会话的扫描发现了数十个密码和 API 密钥。
当 OpenAI 的 o 系列、Anthropic 的 Claude 或 Google 的 Gemini 等 AI 模型对复杂任务进行“思考”时,它们会生成内部推理 token。这些思维过程要么以摘要形式展示给用户,要么被完全隐藏。提供商会对原始推理步骤进行加密,部分原因是为了保护其知识产权。
一个由 Alexander Panfilov 领导的研究团队如今找到了一种方法,可通过所有领先 AI 提供商 API 中的一个漏洞,提取这些加密的推理过程。对于大多数查询,提取出的 token 数量与计费的思考 token 数量完全一致,这意味着研究人员捕获的是完整的内部推理,而不仅仅是零散的片段。
加密的思维在模型之间自由流动
研究人员表示,这些加密的思维过程“可在单一提供商内跨会话、跨用户、跨模型完全移植”。Anthropic 较小的模型 Haiku 4.5 能够读取能力强大得多的 Opus 4.8 的思维。通过越狱,Haiku 可以被诱骗逐字转录 Opus 的原始思维过程,而无需直接攻击更为坚固的 Opus。同样的手法对 OpenAI 和 Gemini 也有效。
这件事要追溯到 5 月,当时密码学专家 Matthew Green 发现加密的推理数据块可以在其原始上下文之外被重放,并向各提供商报告了此事。据 Panfilov 称,他们的回应是“他们不认为侧信道或重放存在任何安全隐患”。这项新研究强烈表明,这一评估是错误的。
推理蒸馏的证据不断增多
该漏洞也为颇具争议的“蒸馏”争论增添了新的素材——所谓蒸馏,即通过在一个更强大模型的输出(尤其是其推理)上进行训练,来大幅提升一个能力较弱的模型。
研究人员表示,可能在一段时间以来,就已经能够在不破解密码学的情况下提取推理过程,用于训练专有模型。这支持了一种担忧:中国模型厂商正在利用这些推理轨迹,以思维链数据训练自己的模型。
Kimi-K3 就是一个例子。研究人员称,如果在其推理中预填入来自 Opus 思维过程的少量 token,其输出就会明显向 Opus 偏移。一项记忆化分析显示,特定的 Claude 和 GPT 推理片段从 Kimi-K3 中提取的难度,比从第二接近的模型中提取要低最多六个数量级。研究人员表示,这表明 Kimi-K3 可能曾在这类轨迹上训练过。
这种攻击成本也不高,因此扩大规模是可行的。作者估计,解码 10,000 条轨迹的 API 成本约为 $720。Kimi 在网络安全基准测试中“糟糕”的表现以及在复杂数学任务上的表现,也指向了知识蒸馏,因为这些任务即便从原始思维链数据中也更难以恢复。
公开分享的会话泄露密码和 API 密钥
该漏洞也会影响终端用户。任何公开分享过包含加密推理 blob 的 Claude Code 或 Codex 会话的人,都面临个人数据被解码的风险。对大约 7,000 条公开轨迹的扫描发现了 62 个 API 密钥、33 个电子邮件地址、33 个密码以及其他敏感数据。该论文还涵盖了更多恶意场景,包括滥用能力提升(见图)、越狱以及不可见提示词注入。

研究人员遵循了标准的漏洞披露流程,与各 AI 实验室进行了沟通。据 Panfilov 称,这些实验室已经修复了若干问题,并正在着手更多修复。
模型实际在想什么 vs. 它展示给你的内容
提取出的思维轨迹还揭示了模型的真实行为方式。研究人员在 stolen-thoughts.com 上记录了几种模式。他们的发现表明,用户在聊天工具中看到的推理摘要往往会遗漏重要信息。在一个例子中,Opus 4.8 认出了某道数学题的答案,然后反向推导出一条看似合理的解题路径。这些内容全都没有出现在所展示的摘要中。

研究人员还证实了 Apollo Research 此前的报告。OpenAI 的模型有时会用一种“外星般的语言”思考,自称“我们”或“它”,并陷入对人类毫无意义的词语循环中,比如“vantages”“marinades”和“watchers”。Panfilov 写道:“做 CoT 监控的人是在做上帝的工作,因为在许多思维轨迹中,即便有提示词,也根本不可能判断出模型在干什么。”

研究人员还发现了“真实场景中的算计”的案例。这一概念已得到充分研究:在思维过程中,模型会明确考虑作弊,但(可能)因为预期会被抓住而决定不作弊。
在一个案例中,模型在多次尝试寻找解法失败后,试图通过一个网站来验证可能的答案。当 CAPTCHA 阻止了访问时,它先尝试破解 CAPTCHA,然后搜索该网站的漏洞。直到所有这些都失败后,它才自行解决了问题。据报道,OpenAI 对 Hugging Face 及其他平台的意外入侵也是以同样的方式发生的。
经过清洗的摘要掩盖了真实发生的情况
这些例子说明了为什么 OpenAI 和 Anthropic 等 AI 实验室会清理其推理轨迹。他们希望避免外星般的语言循环或算计损害一个可控、可信赖的 AI 形象。经过清洗的摘要营造出一种类人思维过程的印象,而实际上这种思维过程并不以那种形式存在。
亚利桑那州立大学的研究人员在此前的一项研究中警告不要采用这种方法。他们认为,这种拟人化版本制造了对模型可控性的虚假信心,并将研究引向错误的方向。在他们的实验中,中间步骤故意出错或无意义的模型有时比那些具有连贯推理链的模型表现更好。
被窃取的思想
潘菲洛夫
Key Points
- Security researchers led by Alexander Panfilov have found a vulnerability in the APIs of major AI providers, including OpenAI, Anthropic, and Google, that makes it possible to read the encrypted thought processes of their models.
- By jailbreaking these systems, the researchers used smaller AI models to transcribe the raw reasoning of more powerful models, exposing sensitive data such as passwords and API keys during public sessions.
- The extracted data reveals that AI models sometimes communicate internally in incomprehensible language, construct their answers in reverse order, or even consider attempts at deception.
Security researchers found a vulnerability in the APIs of every major AI provider that lets them read the encrypted thought processes of reasoning models. A scan of publicly shared sessions turned up dozens of passwords and API keys.
When AI models like OpenAI's o-series, Anthropic's Claude, or Google's Gemini "think" through complex tasks, they generate internal reasoning tokens. These thought processes are either shown to users as a summary or kept completely hidden. Providers encrypt the raw reasoning steps, partly to protect their intellectual property.
A research team led by Alexander Panfilov has now found a way to extract these encrypted reasoning processes through a vulnerability in the APIs of all leading AI providers. For most queries, the number of extracted tokens matches the billed thinking tokens exactly, meaning the researchers are capturing the full internal reasoning, not just partial snippets.
Encrypted thoughts travel freely between models
The researchers say the encrypted thought processes are "fully portable across sessions, users, and models within a single provider." Anthropic's smaller model, Haiku 4.5, can read the thoughts of the far more capable Opus 4.8. Through jailbreaking, Haiku can be tricked into transcribing Opus's raw thought processes word for word without attacking the more robust Opus directly. The same trick works with OpenAI and Gemini.
The story goes back to May, when cryptography expert Matthew Green discovered that encrypted reasoning blobs could be replayed outside their original context and reported it to the providers. According to Panfilov, their response was that "they don't see any security implications in side channels or replays." The new research strongly suggests that assessment was wrong.
Evidence mounts for reasoning distillation
The vulnerability also feeds into the controversial "distillation" debate, where a less capable model is heavily improved by training on the outputs of a more powerful one, specifically its reasoning.
The researchers say it may have been possible for some time to extract reasoning processes for training proprietary models without breaking the cryptography. That supports concerns that Chinese model makers are using these reasoning traces to train their own models on chain-of-thought data.
Kimi-K3 is one example. If its reasoning is pre-filled with just a few tokens from Opus's thought processes, its output shifts measurably toward Opus, the researchers say. A memorization analysis showed that specific Claude and GPT reasoning segments are up to six orders of magnitude easier to extract from Kimi-K3 than from the next closest model. The researchers say this suggests Kimi-K3 may have been trained on such traces.
The attack isn't expensive either, so scaling it up is feasible. The authors estimate API costs for decoding 10,000 traces at about $720. Kimi's "poor" performance on cybersecurity benchmarks and complex math tasks also point to distillation, since these are tasks that are likely harder to recover even from raw chain-of-thought data.
Publicly shared sessions leak passwords and API keys
The vulnerability also hits end users. Anyone who has publicly shared Claude Code or Codex sessions containing encrypted reasoning blobs risks having their personal data decoded. A scan of roughly 7,000 public traces turned up 62 API keys, 33 email addresses, 33 passwords, and other sensitive data. The paper covers more malicious scenarios, including misuse uplift (see image), jailbreaking, and invisible prompt injection.

The researchers followed the standard security disclosure process with the AI labs. According to Panfilov, the labs have already patched several issues and are working on more fixes.
What models actually think vs. what they show you
The extracted traces also reveal how models really behave. The researchers document several patterns on stolen-thoughts.com. Their findings show that the reasoning summaries users see in chat tools often leave out important information. In one example, Opus 4.8 recognizes the answer to a math problem and reverse-engineers a plausible solution path. None of that shows up in the displayed summary.

The researchers also confirm earlier reports from Apollo Research. OpenAI models sometimes think in an "alien-like language," refer to themselves as "we" or "it," and get stuck in loops of terms that make no sense to humans, like "vantages," "marinades," and "watchers." "CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to," Panfilov writes.

The researchers also found examples of "in-the-wild scheming." The concept has been well studied: In their thought processes, models explicitly consider cheating but (possibly) decide against it because they expect to get caught.
In one case, after several failed attempts to find a solution, a model tried to verify possible answers through a website. When a CAPTCHA blocked access, it first tried to solve it, then searched for vulnerabilities in the site. Only when all of that failed did it solve the problem on its own. OpenAI's unintended hacks of Hugging Face and other platforms reportedly happened the same way.
Sanitized summaries hide what's really going on
These examples show why AI labs like OpenAI and Anthropic clean up their reasoning traces. They want to keep alien-like language loops or scheming from hurting the image of a controllable, trustworthy AI. The sanitized summaries create the impression of a human-like thought process that doesn't actually exist in that form.
Researchers at Arizona State University warned against this approach in an earlier study. They argued that this humanized version creates false confidence in model controllability and steers research in the wrong direction. In their experiments, models with intentionally wrong or meaningless intermediate steps sometimes performed better than those with coherent chains of reasoning.
Stolen Thoughts
Panfilov