OpenAI 称 Astra 是其首个达到 Critical 网络安全能力阈值的模型

Rohan Paul · @rohanpaul_ai · X·2026-09-02 10:12·20天前
AI 导读

OpenAI 宣布 Astra 在其 Preparedness Framework 下达到 Critical 网络安全能力阈值,是首个达到该水平的模型,并预告了评估方式与安全措施。

Rohan Paul@rohanpaul_ai
69AI 编辑部评分,满分 100

OpenAI 称 Astra 是其首个达到 Critical 网络安全能力阈值的模型

2026-09-02 10:12· 20天前
AI 导读

OpenAI 宣布 Astra 在其 Preparedness Framework 下达到 Critical 网络安全能力阈值,是首个达到该水平的模型,并预告了评估方式与安全措施。

OpenAI says Astra is its first model to reach the Critical cybersecurity capability threshold.

Under its Preparedness Framework, that means Astra can, with the right tools and access, find unknown flaws and develop exploits across hardened systems without step-by-step human guidance.

Hence, OpenAI now says Astra will launch with additional chain-of-thought monitoring, while classifiers can automatically stop potentially unauthorized actions.

Astra reached roughly 39% exploit success at ~75K output tokens, while GPT-5.6 Sol is only around 1% there and needs nearly 140K tokens to reach ~12%

i.e. Astra is dramatically more capable and token-efficient at exploit development on this internal benchmark.

Expert assessments went further: Astra escaped a browser sandbox, executed host commands, and escalated an unprivileged operating-system user to root.

OpenAIAs we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible. Astra represents a significant advance in cybersecu...

来源:Rohan Paul· x.com