跳到正文
原文
The Decoder:AI News· Maximilian Schreiner·· 3 小时前精选AI 评分77

OpenAI 称拦截蒸馏窃取攻击,但研究者称同样手法在 Azure 上仍可窃取 GPT-6 Astra 等模型的推理内容

OpenAI says it stopped a campaign to steal its models' reasoning, but the trick still worked on Azure

AI 导读

OpenAI 称 7 月拦截了一起针对其模型推理链的蒸馏窃取活动,7 月 24 至 25 日出现来自超 4000 用户的 16000 次请求,关联账号超 15000 个,OpenAI 将其与 Moonshot AI 相关人员联系起来,并于 7 月 28 日关停。

推荐理由

原文把 OpenAI 的拦截行动与研究团队的复测放在一起看,读者可以了解同一模型在不同云平台防护不一致的问题。

正文 · 原文

In distillation, one model learns from another model's full output. That includes the stronger model's complete chains of thought, not just its answers. Those intermediate steps are what make the output so valuable. OpenAI says they can contain information that's deliberately kept out of the final answer, and they can help others recreate a model's capabilities.

According to OpenAI, the activity began at low volume on July 1. On July 24 and 25, it spiked to 16,000 requests from more than 4,000 users, all relying on a typical extraction pattern. A closer look uncovered a network of more than 15,000 accounts with related patterns, which OpenAI says it had fully shut down by July 28. A footnote clarifies that these were attempted extractions, not necessarily successful ones.

OpenAI links a core group behind the activity to people associated with Moonshot AI, the company that makes the Kimi language model. The company says it's unclear whether all the actors it observed trace back to a single source. Anthropic recently reported similar attempts by Chinese AI companies.

A cheap model can unlock an expensive model's hidden thoughts

According to OpenAI, the attackers copied encrypted reasoning from one conversation and then asked a model in a separate conversation to decrypt it and write it out.

Researcher Joachim Schaeffer and his team had already shown how this works in a paper. AI providers send reasoning back to customers only as encrypted data packets, which customers pass along with follow-up requests. Because those packets are encrypted with shared keys, the researchers found, they can be moved between sessions, between users, and even between different models from the same provider. That lets a weaker, cheaper model from the same family act as a "decryption oracle" that prints the stronger model's hidden thoughts word for word.

OpenAI credits the researchers by name. The company says it confirmed that the attack paths they reported were real and that their findings helped it roll out countermeasures faster.

OpenAI says it has since banned fraudulent accounts, tightened sign-ups, and closed the hole that let people reuse and read encrypted reasoning that didn't belong to them. It also now screens streamed outputs and holds them back if they might reveal reasoning. OpenAI says it shared what it learned through the Frontier Model Forum and government channels, because the problem isn't limited to its own models.

Locking down the API doesn't help if the cloud stays open

The story doesn't end there. On the same day, the researchers published an update to their study. "We stole reasoning. Again," Schaeffer wrote on X. Securing your own API, he argues, doesn't secure the wider ecosystem of cloud providers that also sell access to the models.

When the team tested again on September 13, the attack was blocked on OpenAI's and Anthropic's own APIs. On Microsoft Azure, it worked against every OpenAI model they tried, including the new GPT-6 Astra, and against Anthropic models up to Sonnet 5. A single attempt was enough to pull out the reasoning verbatim. "Same models, but different protections depending on which platform serves them," Schaeffer said.

There's also a second, even simpler method, which developer Can Bölük demonstrated publicly. The model gets a virtual notepad as a tool and is told to write its reasoning there, and the user can then read whatever it wrote. The researchers say this worked on every OpenAI model, as well as on Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5, and Fable 5.1 didn't reveal their reasoning. The output closely resembled what the decryption attack produced, and the researchers believe it would likely be just as useful for distillation.

The researchers describe the fixes so far as piecemeal and superficial. Many rely on brittle matching of specific request patterns, and some reached the cloud platforms only days later. GPT-6 Astra launched on third-party platforms without any of the protections in place. According to the researchers' timeline, OpenAI didn't add safeguards to the Azure endpoint until September 27. For Anthropic models, the reported extraction could no longer be reproduced on Azure starting September 28.

Schaeffer argues that patches have to cover every type of attack and every cloud that hosts the models. Otherwise, attackers can simply pick the route with the weakest defenses. The paper goes a step further, arguing that cloud providers that don't enforce equivalent protections shouldn't be allowed to serve reasoning models at all. If they do, the researchers write, open backdoors would effectively let people sidestep export controls at the API level. OpenAI acknowledges that models hosted by partners need the same protection as its own services and says the work is not finished.

We previously covered the research team's original discovery of the vulnerability. Our latest AI Radar newsletter on Chinese language models takes a deeper look at distillation. OpenAI expects these attempts to grow more sophisticated as leading models improve and more players hunt for cheap ways to copy what they can do.

来源:The Decoder:AI News · the-decoder.com