Rohan Paul · @rohanpaul_ai · X·2026-09-11 12:42·2小时前
AI 导读

Anthropic 最新的“AI 滥用”报告称,知识蒸馏可以提升通用推理能力,足以将危险能力扩展到训练对话所涵盖范围之外。 报告还称,Claude 的安全防护无法通过未经授权的蒸馏转移,但未对这些说法提供量化评估。

Rohan Paul@rohanpaul_ai
45AI 编辑部评分,满分 100
2026-09-11 12:42· 2小时前
AI 导读

Anthropic 最新的“AI 滥用”报告称,知识蒸馏可以提升通用推理能力,足以将危险能力扩展到训练对话所涵盖范围之外。 报告还称,Claude 的安全防护无法通过未经授权的蒸馏转移,但未对这些说法提供量化评估。

Anthropic latest "misuse of AI" report claims distillation can improve general reasoning enough to increase dangerous capabilities beyond the subjects covered in the training conversations.

It also says Claude’s safeguards do not transfer through unauthorized distillation, but provides no quantified evaluation for these claims.

来源:Rohan Paul· x.com