Thomas Wolf · @Thom_Wolf · X·2026-08-29 20:42·23天前
AI 导读

Hugging Face 联创 Thomas Wolf 指出,长期看开源与闭源模型面临的安全挑战完全相同,都需在根本行为层面进行稳健、全面且核心的对齐。他认为沙箱、护栏等外部限制手段无法带来廉价安全,唯一出路是让模型自身不想做坏事。

Thomas Wolf@Thom_Wolf
39AI 编辑部评分,满分 100
2026-08-29 20:42· 23天前
AI 导读

Hugging Face 联创 Thomas Wolf 指出,长期看开源与闭源模型面临的安全挑战完全相同,都需在根本行为层面进行稳健、全面且核心的对齐。他认为沙箱、护栏等外部限制手段无法带来廉价安全,唯一出路是让模型自身不想做坏事。

Most people haven’t updated their priors yet, but over the long run, safety challenges are exactly the same for open-source and closed-source models.

You need to align models at a fundamental behavioral level and ensure that this alignment is robust, comprehensive, and core to the model’s behavior.

In the long term, no amount of sandboxing, guardrailing, manifold-limited alignment, or cherry-on-top training will buy you cheap safety.

roonif you think we can contain these things through human ingenuity you’re going to have a bad time in the long run the only recourse you have is to make them not ...