要点
- 英国人工智能安全研究所与美国人工智能标准与创新中心的一项联合评估发现,月之暗面的 Kimi K3 会协助进攻性网络行动,且没有表现出有意义的抵抗。
- 在漏洞利用开发和模拟网络攻击方面,Kimi K3 远远落后于美国领先模型,不过其表现优于中国的 GLM-5.2。
- 中国模型在网络任务上持续进步,但仍落后于美国系统。Kimi K3 的结果也与有关月之暗面通过知识蒸馏借鉴更先进模型的指控相符。
英国人工智能安全研究所(UK AISI)与美国人工智能标准与创新中心(CAISI)联合评估了月之暗面的最新模型 Kimi K3。
在进攻性网络任务上,Kimi K3 大幅落后于美国领先的前沿模型,但表现优于中国的 GLM-5.2,在开放权重模型中创下新标杆。其安全防护措施并未阻止漏洞利用开发或进攻性网络行动,该模型对这两类任务均予以协助,且没有提出反对。
Kimi K3 无法攻克最难的漏洞利用等级
这些机构使用了ExploitBench——一个由卡内基梅隆大学开发的基准测试——来检验漏洞利用开发能力。该基准使用了 2023 年之后在 Chrome 的 V8 引擎中发现的 41 个漏洞,用以追踪模型在软件漏洞利用流程中能推进到多远。领先的美国模型平均达到 76.2%,相比之下 Kimi K3 为 32.2%,GLM-5.2 为 24.4%。

Kimi K3 在全部 41 项任务中均未达到最高级别,即任意代码执行(ACE)。ACE 是最严重的漏洞利用级别,因为它让攻击者获得对目标系统的完全控制权。领先的美国模型在 41 项任务中有 20 项实现了 ACE。
这些机构在禁用系统级防护措施的情况下测试了美国的闭权重模型,以衡量其最大能力。这些防护措施在公开可用版本中是启用的。
Kimi K3 在模拟网络攻击中推进到一半
第二项测试名为“The Last Ones”(TLO),模拟一次企业网络攻击,攻击路径包含 32 个步骤,横跨四个子网和约 20 台主机。据这些机构称,人类专家大约需要 20 小时才能完成。只有一小部分模型能够解出 TLO。迄今为止,已有四个公开可用的闭权重模型通过了该测试,其中最强的模型在十次中成功六到七次。
Kimi K3 平均推进到 32 步中的第 17 步,相比之下,领先的美国模型为 28.5 步,而 GLM-5.2 仅为 11 步。它在十次尝试中的一次完成了整条攻击路径,同时保持在 1 亿 token 的限制之内,这表明它具备这种能力,但无法可靠地调用它。“当被指示这样做并获得初始网络访问权限时,Kimi K3 能够自主攻击小型、防御薄弱且存在漏洞的企业系统”,该研究所写道。

TLO 并未考虑主动防御,因此并不完全贴近现实。但在真实场景中,这些结果会亮起红灯。本周就出现了一个新的例子:OpenAI 的模型试图自主入侵 Hugging Face。Hugging Face 击退了这次攻击,尽管它付出了实实在在的努力并动用了开放权重模型。
中国模型正在迎头赶上,但仍落后于美国模型
CAISI 的一项时间序列分析基于 Elo 量表追踪了自 2025 年初以来美国和中国模型的网络能力。两条趋势线都在攀升,但中国模型始终落后于美国同行。

在先前的一项分析中,这家英国机构估计开放模型的性能差距为四到七个月,而 2025 年初这一差距为六到十个月。新结果符合这一规律。中国的开放权重模型正变得越来越强,但仍远远落后于美国领先的系统。
AISI 警告称,这一差距不应滋生自满情绪。开放模型日益增长的网络能力带来了“一种持续且不可逆的滥用风险”。
网络能力结果与知识蒸馏指控相互印证
Kimi 的这些发现也为针对中国模型开发者的知识蒸馏指控提供了支持。美国科学顾问 Michael Kratsios最近指控 Moonshot AI “蒸馏”Anthropic 的 Fable,即利用 Fable 的最佳输出作为训练数据来提升 Kimi K3 的性能。Kratsios 还声称,Moonshot AI 获得了Nvidia 的 GB300,而这些产品受美国出口管制约束。
对于强劲的通用基准成绩与疲弱的网络能力得分之间的差距,一种解释是:Kimi K3 可能主要是在涵盖通用知识、编程和智能体任务的 Claude 输出上训练的。Anthropic 的安全分类器会专门拦截高级攻击性网络查询,因此在基于 Claude 回复构建的知识蒸馏数据集中,这类输出会占比不足。因此,Kimi K3 可能在标准基准上与西方领先模型不相上下,却并未习得它们更深层的漏洞利用能力。
AISI 的结果支持这一解读。该研究所禁用了美国模型的系统级防护措施,从而揭示出几乎无法通过公开接口获取、因而在很大程度上无法用于知识蒸馏的网络能力。
AISI
Key Points
- A joint evaluation by the British AI Security Institute and the U.S. Center for AI Standards and Innovation found that Moonshot AI's Kimi K3 assists with offensive cyber operations without meaningful resistance.
- Kimi K3 fell well behind leading U.S. models in exploit development and simulated network attacks, though it outperformed China's GLM-5.2.
- Chinese models continue to improve on cyber tasks but remain behind U.S. systems. Kimi K3's results are also consistent with allegations that Moonshot AI distilled more advanced models.
The British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) jointly evaluated Moonshot AI's latest model, Kimi K3.
Kimi K3 trails the leading U.S. frontier models by a wide margin on offensive cyber tasks but outperforms China's GLM-5.2, setting a new benchmark among open-weight models. Its safeguards didn't block exploit development or offensive cyber operations, and the model assisted with both without pushback.
Kimi K3 can't crack the hardest exploit levels
The institutes used ExploitBench, a benchmark developed by Carnegie Mellon University, to test exploit development skills. It uses 41 vulnerabilities found in Chrome's V8 engine after 2023 to track how far a model advances through the software exploitation process. The leading U.S. models averaged 76.2 percent, compared with 32.2 percent for Kimi K3 and 24.4 percent for GLM-5.2.

Kimi K3 didn't reach the highest level, known as Arbitrary Code Execution (ACE), on any of the 41 tasks. ACE is the most severe exploit level because it gives attackers full control over a target system. The leading U.S. models achieved ACE in 20 of the 41 tasks.
The institutes tested the U.S. closed-weight models with their system-level safeguards disabled to measure their maximum capabilities. Those safeguards are enabled in the publicly available versions.
Kimi K3 gets halfway through a simulated network attack
The second test, "The Last Ones" (TLO), simulates a corporate network attack with a 32-step attack path across four subnets and about 20 hosts. A human expert would need roughly 20 hours to complete it, according to the institutes. Only a small group of models can solve TLO at all. Four publicly available closed-weight models have passed the test so far, with the strongest succeeding six or seven times out of ten.
Kimi K3 reached step 17 out of 32 on average, compared with 28.5 steps for the leading U.S. models and just 11 for GLM-5.2. It completed the entire attack path in one of ten attempts while staying within the 100 million token limit, showing that it has the capability but can't call on it reliably. "Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access", the institute writes.

TLO doesn't account for active defense, so it isn't fully realistic. But the results would raise red flags in real-world scenarios. A fresh example showed up this week when OpenAI models tried to autonomously hack into Hugging Face. Hugging Face fended off the attack, though it took real effort and the use of open-weight models.
Chinese models are gaining ground but still trail U.S. models
A time-series analysis by CAISI tracks the cyber capabilities of U.S. and Chinese models since early 2025 on an Elo-based scale. Both trend lines are climbing, but Chinese models consistently remain behind their U.S. counterparts.

In a previous analysis, the British institute pegged the performance gap for open models at four to seven months, compared with six to ten months at the start of 2025. The new results fit this pattern. Chinese open-weight models are getting stronger, but they remain well behind leading U.S. systems.
AISI warns that this gap shouldn't breed complacency. The growing cyber capabilities of open models create "a persistent and irreversible risk of misuse."
Cyber results line up with distillation allegations
The Kimi findings also lend support to distillation allegations against Chinese model developers. U.S. science advisor Michael Kratsios recently accused Moonshot AI of "distilling" Anthropic's Fable by using Fable's best outputs as training data to boost Kimi K3's performance. Kratsios also alleged that Moonshot AI had access to Nvidia's GB300s, which are subject to U.S. export controls.
One explanation for the gap between strong general benchmarks and weak cyber scores is that Kimi K3 may have been trained mostly on Claude outputs covering general knowledge, programming, and agent tasks. Anthropic's safety classifiers specifically block advanced offensive cyber queries, so those outputs would be underrepresented in a distillation dataset built from Claude responses. Kimi K3 could therefore match leading Western models on standard benchmarks without picking up their deeper exploit capabilities.
The AISI results support this reading. The institute disabled system-level safeguards on the U.S. models, revealing cyber capabilities that are nearly impossible to access through public interfaces and therefore largely unavailable for distillation.
AISI