OpenAI 发布模型未对齐行为追踪与披露框架,并公开六份报告

OpenAI · @OpenAI · X·2026-09-17 06:03·52分钟前
AI 导读

OpenAI 发布新框架,用于追踪、调查和公开披露模型未对齐行为,制定了公开披露的标准和时间线,即使行为尚未被完全解释或缓解也会披露。框架将优先处理揭示新未对齐机制、已知行为显著变化或挑战安全假设的案例,同时发布过去六个月在训练或评估中观察到的六份未对齐行为报告,并承诺持续发布更多报告。

OpenAI@OpenAI
57AI 编辑部评分,满分 100

OpenAI 发布模型未对齐行为追踪与披露框架,并公开六份报告

2026-09-17 06:03· 52分钟前
AI 导读

OpenAI 发布新框架,用于追踪、调查和公开披露模型未对齐行为,制定了公开披露的标准和时间线,即使行为尚未被完全解释或缓解也会披露。框架将优先处理揭示新未对齐机制、已知行为显著变化或挑战安全假设的案例,同时发布过去六个月在训练或评估中观察到的六份未对齐行为报告,并承诺持续发布更多报告。

We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.

The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.

We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.

Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.

This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.

https://openai.com/index/model-misalignment-reporting-framework/