一个由数十名网络安全专家组成的团体,其中包括多位业内知名资深人士,向美国政府发布了一封公开信,要求其解除对Anthropic的Fable和Mythos模型的出口管制令。
根据这封公开信,“这一行动从[网络安全]防御者手中夺走了最好的模型”,导致他们现在无法使用这些模型来发现漏洞,并使其软件和产品更加安全。
信中写道:“在我们的对手快速进步之际,毫无充分理由地从防御者手中夺走最强大的能力,这是危险的。”
据Anthropic称,上周五,美国政府以国家安全为由,命令Anthropic限制Fable和Mythos的出口,但未说明该命令背后的具体原因。作为回应,该公司暂停了全球所有用户对这些模型的访问权限。
截至本文撰写时,这封公开信已有76位网络安全专家签署,其中包括:前Facebook首席安全官Alex Stamos;漏洞赏金平台Bugcrowd创始人Casey Ellis;著名密码学家、前苹果安全设计与架构经理Jon Callas;计算机科学家Paul Vixie;Block前应用安全工程主管Dino Dai Zovi;Luta Security创始人Katie Mossouris;以及安全意识培训公司SocialProof Security首席执行官Rachel Tobac。
当Mythos于4月以预览版形式推出时,Anthropic声称它在发现安全漏洞方面非常强大,以至于公司需要严格限制访问,以防止恶意黑客或外国对手利用它在互联网上造成破坏。实际上,这意味着Anthropic最初只给了大约50家公司Mythos的初始访问权限,最近才将该群体扩大到包括15个国家的约150个组织。
上周,Anthropic 发布了 Fable,这是 Mythos 的一个公开版本。该公司表示,该版本设置了严格的护栏,以阻止其在生物学、化学和网络安全领域的应用,并防止他人对该模型进行知识蒸馏以复制它。Fable 上的护栏极其严格,以至于许多网络安全专家发现,它基本上会阻止任何与网络安全相关的提示词。
Anthropic 表示,白宫的出口管制令可能是基于一份报告,该报告称存在一种绕过(即所谓的越狱)Fable 的方法,以解锁其强大的 Mythos 级别能力。
据公开信的签署人之一 Katie Moussouris 称,该方法由亚马逊研究人员在一篇尚未公开的论文中进行了演示,但她已经审阅过该论文。
但 Moussouris 在一篇博客文章中表示,该论文实际上并未演示真正的越狱。相反,她写道,研究人员只是要求 Fable 修复那些存在公开已知漏洞的开源代码,并“故意植入漏洞”,而在此之前,该模型最初拒绝“审查代码中的安全问题”。
“论文中描述的行为无法得到有意义的修复,任何尝试都只会削弱模型在防御方面的能力,”Moussouris 写道。“防御方需要能够要求 AI 修复文件中的错误,解释修复为何重要,并编写测试来确认补丁有效。这不是绕过护栏。这是 AI 模型能为防御性安全所做的最有价值的事情:执行防御方每天都在进行的发现、修复和测试循环。”
Moussouris 的批评在公开信中得到了呼应,信中还表示,专家小组认为亚马逊论文中的方法“可以在 OpenAI 的 GPT-5.5、Anthropic 自己公开发布的 Claude Opus 4.8 和 Sonnet,甚至像 Kimi 2.7 这样的中国模型上被复制。”
这封信还要求通过“民主的规则制定过程”来制定透明且公平执行的法规,这些法规应基于行业和学术专家进行的科学研究,并且“仅在确保美国公众安全所必需的最低限度内使用”。
A group made up of dozens of cybersecurity experts, including several well-known veterans of the industry, published an open letter to the U.S. government asking it to lift the export control order on Anthropic’s Fable and Mythos models.
According to the open letter, “this action has taken the best models away from [cybersecurity] defenders” who now can’t use the models to find vulnerabilities and make their software and products more secure.
“To pull the best capabilities away from defenders without a good reason when our adversaries are rapidly advancing is dangerous,” read the letter.
On Friday, the U.S. government ordered Anthropic to limit the export of Fable and Mythos citing national security concerns, without explaining the specific reasons behind the order, according to Anthropic. In response, the company suspended access to the models to all users worldwide.
As of this writing, the letter is signed by 76 cybersecurity experts, including: former Facebook chief of security Alex Stamos; Casey Ellis, the founder bug bounty platform Bugcrowd; famed cryptographer and former Apple security design and architecture manager Jon Callas; computer scientist Paul Vixie; Dino Dai Zovi, the former head of applied security engineering at Block; Katie Mossouris, the founder of Luta Security; and Rachel Tobac, the CEO of the security awareness training firm SocialProof Security.
When Mythos launched as a preview in April, Anthropic claimed it was so powerful at finding security vulnerabilities that the company needed to tightly restrict access to prevent malicious hackers or foreign adversaries from using it to cause havoc on the internet. In practice, that meant Anthropic gave around 50 companies initial access to Mythos, recently expanding that group to include around 150 organizations in 15 countries.
Last week, Anthropic released Fable, a public version of Mythos that the company said had strict guardrails to block its use in the fields of biology, chemistry, and cybersecurity, as well as to stop others from distilling the model in order to recreate it. The guardrails on Fable were so strict that many cybersecurity experts found that it stopped essentially any prompts related to cybersecurity.
Anthropic said that the White House export control order may have been based on a report that there was a method to bypass — or so-called jailbreaking — Fable to unlock its powerful Mythos-level capabilities.
According to Katie Moussouris, one of the signatories of the open letter, the method was demonstrated by Amazon researchers in a paper that is not public, but that she has reviewed.
But Moussouris said in a blog post that the paper did not actually demonstrate a real jailbreak. Instead, she wrote, the researchers simply asked Fable to fix open source code with public and known vulnerabilities along with “deliberately planted vulnerabilities,” after the model initially refused to “review the code for security issues.”
“The behavior described in the paper cannot meaningfully be fixed, and any attempt would only weaken the model for defense,” Moussouris wrote. “Defenders need to be able to ask AI to fix the bugs in a file, explain why the fix matters, and write tests that confirm the patch works. That is not a guardrail bypass. It is the most valuable thing an AI model can do for defensive security: executing the find, fix, and test loop defenders run every day.”
Moussouris’ critique was echoed in the open letter, which also said that the group of experts believe the method in the Amazon paper “can be replicated” on OpenAI’s GPT-5.5, on Anthropic’s own publicly-available Claude Opus 4.8 and Sonnet, “and even Chinese models like Kimi 2.7.”
The letter also asked for transparently and fairly enforced regulations created by “a democratic rule-making process” that are based on scientific research done by industry and academic experts, and “used only to the minimal extent necessary to ensure the safety of the American public.”