Hugging Face 联手 Baseten、Goodfire 推进开源模型安全研究

Thomas Wolf · @Thom_Wolf · X·2026-09-17 04:53·2小时前
AI 导读

Hugging Face 联合创始人 Thomas Wolf 宣布与 Baseten 旗下 Baselabs 及 GoodfireAI 合作,推进开源模型的安全与可解释性研究。合作方将把安全研究直接构建进 Baseten 的推理基础设施,包括训练模型遵循明确策略、运行时检测故障,并将信号接入可干预的控制机制,目标向所有客户开放。Wolf 强调开源与安全部署并不矛盾,正如 Linux 所证明的那样。

Thomas Wolf@Thom_Wolf
37AI 编辑部评分,满分 100

Hugging Face 联手 Baseten、Goodfire 推进开源模型安全研究

2026-09-17 04:53· 2小时前
AI 导读

Hugging Face 联合创始人 Thomas Wolf 宣布与 Baseten 旗下 Baselabs 及 GoodfireAI 合作,推进开源模型的安全与可解释性研究。合作方将把安全研究直接构建进 Baseten 的推理基础设施,包括训练模型遵循明确策略、运行时检测故障,并将信号接入可干预的控制机制,目标向所有客户开放。Wolf 强调开源与安全部署并不矛盾,正如 Linux 所证明的那样。

Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpretability for open models.

As open-source models catch up in performance to frontier models, the community at large have a great opportunity to establish a common practice for effective safety, security and interpretability research.

Some past discussions confused “open” with “unmonitored”. On the contrary, an open model provides much more tooling and visibility for safety research from a broader audience, which has led to a lot of the safety and security techniques we use today.

Linux is a great example of this: open-source and secure deployment are not only compatible but heavily intertwined, as a properly secure system needs an extensive feedback loop of finding and fixing vulnerabilities.

We're excited to share more soon.

Charlie O'NeillSafety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future...

来源:Thomas Wolf· x.com