Rohan Paul · @rohanpaul_ai · X·2026-09-17 06:40·2小时前
AI 导读

OpenAI 发布追踪、调查和披露模型对齐问题的新框架,设定公开披露的标准和时间线,即使尚未完全解释或修复该行为也会公布。框架优先披露揭示新失当机制、已知问题显著变化或挑战现有安全假设的案例,并同步发布过去六个月训练或评估中观察到的六份失当行为报告,后续将根据公开反馈持续完善并发布更多报告。

Rohan Paul@rohanpaul_ai
65AI 编辑部评分,满分 100
2026-09-17 06:40· 2小时前
AI 导读

OpenAI 发布追踪、调查和披露模型对齐问题的新框架,设定公开披露的标准和时间线,即使尚未完全解释或修复该行为也会公布。框架优先披露揭示新失当机制、已知问题显著变化或挑战现有安全假设的案例,并同步发布过去六个月训练或评估中观察到的六份失当行为报告,后续将根据公开反馈持续完善并发布更多报告。

So OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior.

So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports.

They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.

OpenAIWe're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines...

来源:Rohan Paul· x.com