# OpenAI 发布模型对齐问题披露框架，并公开六份失当行为报告

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-17 06:40
- AIHOT 分数：65
- AIHOT 链接：https://aihot.news/items/cmu4oro0t08eqrodc194qfslw
- 原文链接：https://x.com/rohanpaul_ai/status/2100354109664813407

## AI 摘要

OpenAI 发布追踪、调查和披露模型对齐问题的新框架，设定公开披露的标准和时间线，即使尚未完全解释或修复该行为也会公布。框架优先披露揭示新失当机制、已知问题显著变化或挑战现有安全假设的案例，并同步发布过去六个月训练或评估中观察到的六份失当行为报告，后续将根据公开反馈持续完善并发布更多报告。

## 正文

So OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior.

So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports.

They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.

### 引用推文

> OpenAI：We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines...
