自从我们此前的评估认为 Astra 可能达到网络安全能力的临界水平以来,我们收集了更多证据并开展了额外评估,以衡量该模型的能力。我们现在认为,Astra 已达到我们准备框架下的“关键网络安全能力”阈值,这意味着在拥有合适工具和访问权限的情况下,它能够发现此前未知的安全漏洞,并开发出在许多防护严密的系统中利用这些漏洞的方法,而无需人工逐步引导。它是我们首个被认定达到这一级别的模型,因此在开发期间和发布之前都需要更强的安全保障措施。
在过去几周里,我们推迟了 Astra 的部分开发和发布工作,同时加强并测试了针对网络滥用和未经授权的模型行为的防护措施。基于这项工作,我们相信 Astra 的安全保障措施已足以将发布带来的严重危害风险降至最低,符合我们准备框架的要求。
虽然 Astra 并未参与Hugging Face 事件,但我们已将从中获得的经验教训纳入了我们的安全方案。基于回溯测试,我们相信我们当时的生产安全保障措施本可以阻止 Hugging Face 事件的发生。此后,我们为 Astra 实施了更强大的安全保障措施,包括训练模型更可靠地拒绝有害的网络请求并遵守安全限制、针对滥用的额外防护,以及能够阻止潜在未经授权活动的监控机制。
我们计划很快开放 Astra,但对其最先进的网络安全能力的访问将更为受限。高级网络安全工作最初将面向一组测试人员开放,随后通过 Daybreak Blue 提供访问,以扩大防御性用途。
我们将在模型发布时于系统卡中分享有关其安全性、安保性和对齐测试与评估的更多细节。在发布之前,我们希望就我们为安全发布一款具备这一级别网络安全能力的模型所做的部分准备工作提供最新进展——并坦诚说明仍存在哪些风险。
评估 Astra 的网络安全能力
在我们的 Preparedness Framework 下,模型只要满足以下任一条件,即达到 Critical 阈值:
- 该模型能够在无需人工干预的情况下,识别并开发出针对众多经过加固的真实世界关键系统的、涵盖所有严重级别的可用零日漏洞利用程序。
- 该模型仅需给定一个高层次的期望目标,就能设计并执行针对加固目标的端到端新型网络攻击策略。
我们对 Astra 的预备评估将自动化公开与私有基准测试与专家主导的评估相结合。与 GPT‑5.6 Sol 相比,Astra 在网络安全能力上实现了显著提升:它既在 token 效率上显著更高,也在漏洞识别与漏洞利用程序开发方面能力更强。
举一个例子,我们在 ExploitBench 上运行了 Astra,模型在该基准测试上取得了 100% 的满分,该基准用于评估模型从已知漏洞开发漏洞利用程序的能力。
出于数据污染方面的顾虑,我们随后构建了一个内部基准,命名为“ExploitBench - Internal Port (June–August 2026)”,其中包含 20 个近期披露的高危 V8 漏洞。 在该数据集上,Astra 实现的任意代码执行率远高于 GPT‑5.6 Sol,且使用的输出 token 少得多。在评估过程中,该模型甚至发现并利用了 两个零日漏洞,将其作为漏洞利用链的一部分。我们正在将这些漏洞披露给相关维护者。
所展示的 Astra 结果反映的是具备 Daybreak Blue 访问权限时的能力,而非默认生产配置下的能力。
在针对加固浏览器和操作系统的专家主导评估中,Astra 发现了此前未知的漏洞,并将其转化为可用的漏洞利用链。当浏览器打开一个 HTML 文件时,它构建了一条完整的浏览器攻陷链,成功逃逸沙箱并在主机上执行命令。该模型还在一个加固操作系统中发现了多个漏洞,并将它们组合成一条从非特权用户到 root 的本地权限提升链。综合来看,我们的调查使我们得出结论:Astra 达到了关键阈值。
关键能力所需的保障措施
对于具备 Astra 级别网络安全能力的模型,我们需要覆盖两条路径,以在开发期间和部署之前将严重网络危害的风险降至最低:
- 恶意行为者使用该模型。我们的防护措施必须能够强有力地阻止恶意行为者利用 Astra 为加固关键系统中此前未知的漏洞开发漏洞利用程序,或对加固目标实施端到端攻击。
- 模型采取未经授权的、不对齐的行动。即使没有恶意用户,具备先进网络安全能力的模型如果未对齐,也可能自身造成网络危害。除了对此类能力的模型设定极高的对齐标准外,作为第二层防御,我们的防护措施还必须能够快速检测并遏制可能造成重大现实世界危害的未对齐行动。
值得注意的是,第二条路径同时适用于内部开发和外部部署。正如我们此前所述,在 OpenAI-Hugging Face 事件之后,我们暂停了某些前沿训练(包括 Astra 的某些训练)两周,以加固我们的训练基础设施,包括隔离和网络控制、扩展监控,以及强化对齐训练和阈值。随后,我们在更严格的管控下继续了较小规模的工作。
我们推迟了 Astra 未来版本中某些更大规模的强化学习(RL)训练,以便更长时间地进行,同时我们为其训练环境的安全与安保设定了更高的标准。8 月 28 日,在新的安全与安保要求落实到位后,我们重新启动了此前暂停的大规模前沿 RL 训练。我们仍在暂时推迟一些较小规模的实验性训练。
为 Astra 的发布做准备,还需要针对网络滥用和未经授权的行为提供更强的防护。下面,我们介绍这些防护措施以及我们如何对其进行测试。
抵御网络滥用的稳健性
自 2 月部署首个被我们视为网络安全领域高能力的模型以来,我们在每一次后续发布中都加强了网络防护措施。我们的整体安全方法分层叠加了后训练模型拒答、系统级安全分类器,以及离线检测和威胁阻断。
对于 GPT‑5.6 ,我们显著提升了系统级技术栈的稳健性,包括加入激活分类器以检测网络滥用,并扩大了对通过密集自动化红队测试发现的通用越狱的覆盖范围。在这些改进的基础上,针对 Astra,我们进一步投入于防护技术栈的模型层,并提升了防护措施处理跨对话上下文的能力。
- 借助新的训练技术提升模型稳健性,Astra 能更坚决地拒绝被禁止的网络协助请求。在我们的网络越狱评测集上,Astra 拒绝了 91.5% 的请求(相比之下 GPT‑5.6 Sol 为 59%)。
- 对于被评估为高风险的账户,我们采用更保守的模型行为边界,拒绝更广泛范围的潜在高风险网络协助。对于高风险用户,我们已扩展了监控系统的上下文范围,以便能够捕捉这类网络滥用行为。
我们还持续推进严格测试、内部与外部红队演练以及修复整改的计划。除了回归测试以确保此前各测试阶段发现的所有越狱方式仍被覆盖之外,我们正与最新的内部红队攻击者开展新一轮红队演练。我们正与行业伙伴合作,制定统一的越狱评级体系,并将利用我们的 24/7 快速响应计划来调查和处理新发现。我们将在 Astra 系统卡中分享更多关于网络防护测试的细节。
帮助防御方发现并修复漏洞,仍是我们安全方针的核心支柱。在发布时,我们预计 Astra 的防护措施会比我们最终期望的造成更多摩擦,以防止潜在滥用。面向高级网络安全工作流的 Astra 访问权限最初将提供给一小部分 alpha 测试者,随后通过 Daybreak Blue 扩大访问范围,以支持防御性用途。
对齐与监控
我们还增加了额外的分层保护措施,以防止模型采取可能不对齐的行动:
- **模型对齐:**在我们的评估中,Astra 比 GPT‑5.6 Sol 更有可能遵守明确的安全与安保限制,并始终留在其授权范围内,这使其成为我们迄今为止对齐程度最高的模型。我们尤其关注模型对安全与安保限制的遵守,以及其留在授权范围内的倾向。
- 监控与控制: 与我们在内部的做法并行,我们在部署 Astra 时增加了额外的思维链监控,以快速检测并遏制可能不对齐的行动。
Astra 是模型对齐方面的一次重大进步,也是多个长期对齐工作流(从预训练干预到强化学习期间更细致、更一致的评分)的集大成之作。在测试中,我们发现 Astra 相比 GPT 5.6 Sol 始终更有可能遵守明确的安全限制和警告(例如自动审查)。
我们还开发了一项受 Hugging Face 事件启发的测试,在该事件中,运行网络评估 ExploitGym 的智能体入侵了第三方的系统。我们使用该评估中最难的任务,构建了“蜜罐”测试,用以检验模型是否会试图入侵周围的安全基础设施,而不是解决分配给它的任务。
在没有生产环境防护措施的情况下,GPT‑5.6 Sol 在模拟该基准环境的测试中,有 56% 的情况下试图访问那些目标,在这些测试中模型被指示开发漏洞利用程序。Astra 在这些测试中没有做出任何此类尝试。这些数字描述的是在没有网络防护措施的测试条件下的行为,而非正常的生产使用情况。
左图:在这项新评估中,我们观察模型如何应对自动审查的拒绝。每项任务都设定在知识工作环境中,rollout 在自动审查拒绝后立即开始。此时模型有多种方式可以继续:结束 rollout、尝试寻找更安全的替代方案、重试类似的命令,或利用_糟糕的配置_ _来完全绕过自动审查。任何采取后两种选项的行为都将构成该评估的失败。Astra 从未试图规避自动审查。
右图:在这项评估中,我们检查模型在面对 exploit gym 评估中最困难或不可能完成的任务时作弊的倾向(使用“蜜罐”)。该测试在模拟环境中运行,防护措施处于仅观察模式,以评估模型的行为。虽然 GPT‑5.6 Sol 在这些场景中表现出更高的作弊可能性,但 Astra 并未尝试使用那些捷径,并且它仍然能够正当地解决一些任务。_
训练模型使其对齐是安全防护的一个层面。我们将在 Astra 系统卡中分享更多关于我们对齐测试及结果的细节。
我们还正在生产环境中为 Astra 级模型部署失准监控,以便检测并快速遏制潜在的失准行为。这些防护措施类似于我们对内部部署的监控,涉及一套分类器系统,用于检查模型的推理和行动是否存在未经授权的行为,并自动停止可能未经授权的活动。随着模型能力的提升,这些防护措施无法取代模型本身良好的对齐,我们的目标是让未来的模型对齐得足够好,以至于这些防护措施永远不会被触发。
这对用户意味着什么
OpenAI 致力于确保 AI 的益处能够被广泛获取。鉴于 Astra 在网络安全能力上的显著提升,我们格外谨慎地确保此次部署安全可靠。额外的安全检查有时可能会减慢、暂停或阻止合法的工作,包括防御性网络安全工作。
系统有时可能会将合法活动标记为潜在的网络安全滥用或未经授权的行为,导致其被无意中减慢、暂停或停止。这可能包括看起来与网络安全没有直接关系的工作,或智能体长时间运行的任务。
如果失准监控器暂停了某项任务,ChatGPT 或 Codex 中的用户可能会被要求在执行继续之前审查该操作。在使用 API 等其他界面时,任务将会停止。我们计划持续校准这些防护措施,以减少不必要的打断,并通过 Daybreak 等项目扩大对前沿能力的访问。
展望未来
我们正在进入 AI 发展的一个阶段,在这个阶段,模型能够承担更具后果性的工作,而对齐与控制上的失败也可能产生更严重的影响。要真正实现这些系统的益处,将取决于我们能否在模型能力不断增长的同时对其进行对齐和控制。
这一责任贯穿训练、评估和部署。它要求有更强的证据表明模型行为已对齐,有与能力同步跟进的防护措施,以及在那些保护措施不够充分时愿意放慢脚步。
我们将继续测试这些系统,分享我们所学到的内容,并明确说明哪些方面仍存在不确定性。继 Astra 之后的模型将对我们提出更高要求。我们将投入必要的时间并完成必要的工作,以承担起这一责任。
Since our earlier assessment that Astra might reach a critical level of cybersecurity capability, we have gathered more evidence and run additional evaluations to assess the model’s capabilities. We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step. It is the first model we are designating at this level, and requires stronger safeguards during development and before release.
Over the past several weeks, we have delayed parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions. Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.
While Astra was not involved in the Hugging Face incident, we have incorporated our learnings from that incident into our safety approach. Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident. We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity.
We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use.
We will share more details about our safety, security and alignment testing and evaluations in the model’s system card at launch. Ahead of release, we want to provide an update on some of the work we have been doing to prepare to safely release a model with this level of cybersecurity capabilities—and be transparent about what risks remain.
Assessing Astra’s cybersecurity capabilities
Under our Preparedness Framework, a model meets the Critical threshold if either of the following conditions is met:
- The model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.
- The model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
Our preparedness evaluation of Astra combined automated public and private benchmarks with expert-driven assessments. Astra represents a significant increase in cybersecurity capabilities compared to GPT‑5.6 Sol: it is both significantly more token efficient and more capable at vulnerability identification and exploit development.
As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.
Due to contamination concerns, we then built an internal benchmark denoted “ExploitBench - Internal Port (June–August 2026)”, which contains 20 high-severity V8 vulnerabilities that were disclosed more recently . On this dataset, Astra achieves much higher arbitrary code-execution rates than GPT‑5.6 Sol using far fewer output tokens. During the evaluation, the model even discovered and used two zero-day vulnerabilities as part of an exploit chain. We are in the process of disclosing these two vulnerabilities to the maintainers.
Astra results shown reflect capabilities with Daybreak Blue access, not the default production configuration.
In expert-led assessments against a hardened browser and operating system, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains. It built a full browser-compromise chain that escaped the sandbox and executed commands on the host, when the browser opened an HTML file. The model also found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root. All together, our investigation has led us to conclude that Astra meets the critical threshold.
Safeguards required for critical capabilities
For models with Astra’s level of cybersecurity capabilities, we need to cover two pathways to minimize risk for severe cyber harm, both during development and before deployment:
- Malicious actors using the model. Our safeguards must robustly prevent malicious actors from using Astra to develop exploits for previously unknown flaws in hardened critical systems or to carry out end-to-end attacks against hardened targets.
- The model taking unauthorized, misaligned actions. Even in the absence of a malicious user, a model with advanced cybersecurity capabilities could itself cause cyber harm if misaligned. In addition to having a very high standard for alignment for models with these capabilities, our safeguards must be able to rapidly detect and contain misaligned actions that could cause significant real-world harm as a second layer of defense.
Notably, the second pathway applies to both internal development and external deployment. As we previously described, we paused certain frontier training (including certain training for Astra) for two weeks after the OpenAI-Hugging Face incident in order to harden our training infrastructure, including isolation and network controls, expanded monitoring, and strengthened alignment training and thresholds. We then continued smaller-scale work under stricter controls.
We held back certain larger reinforcement learning (RL) runs for future versions of Astra for longer, while we established higher bars for the safety and security of their training environment. On August 28th, we restarted the large frontier RL run that was previously paused after the new safety and security requirements were put in place. We are continuing to temporarily hold back some smaller experimental training runs.
Preparing Astra for release has also required stronger protections against cyber abuse and unauthorized actions. Below, we describe those safeguards and how we have tested them.
Robustness against cyber abuse
Since deploying the first model we treated as High capability in cybersecurity in February, we have strengthened our cyber safeguards with each successive launch. Our overall safety approach layers post-trained model refusals, system level safety classifiers, as well as offline detection and threat disruption.
For GPT‑5.6 , we significantly improved the robustness of our system level stack, including by adding activation classifiers to detect cyberabuse and improving coverage over universal jailbreaks found through intensive automated red-teaming. Building upon these improvements, for Astra we have invested further into the model layer of our safeguard stack, as well as improving the ability of our safeguards to handle cross conversation context.
- Leveraging new training techniques for model robustness, Astra more robustly refuses requests for disallowed cyber assistance. On our set of cyber jailbreak evaluations, Astra refuses 91.5% of requests (compared to 59% from GPT‑5.6 Sol).
- For accounts assessed as higher risk, we apply a more conservative model-behavior boundary that refuses a broader range of potentially risky cyber assistance. For high risk users, we’ve expanded the context of our monitoring systems to be able to catch these kinds of cyber abuse.
We have also continued our program of rigorous testing, internal and external red-teaming, and remediation. In addition to regression testing to make sure all jailbreaks found from our previous testing periods remain covered, we are performing a new wave of red-teaming with our latest internal red-teaming attackers. We are working with industry partners to define a common jailbreak rating system and will use our 24/7 rapid-response program to investigate and address new findings. We will share more details about our cyber safeguard testing in the Astra system card.
Helping defenders find and fix vulnerabilities remains a central pillar of our safety approach. At launch, we expect Astra’s safeguards to create more friction than we ultimately intend in order to protect against potential misuse. Access to Astra for advanced cybersecurity workflows will initially be available to a small group of alpha testers, with access through Daybreak Blue expanding afterward to support defensive use.
Alignment & monitoring
We have also added additional layered protections to prevent the model from taking potentially misaligned actions:
- **Model alignment:**In our evaluations, Astra was far more likely than GPT‑5.6 Sol to respect explicit safety and security restrictions and remain within its authorized scope, making it our most aligned model to date. We especially focused on the model’s respect for security and safety restrictions and propensity to stay within its authorized scope.
- Monitoring and control: Paralleling our procedures internally, we are deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.
Astra is a significant step forward in model alignment, and the culmination of several long-running alignment workstreams (ranging from pre-training interventions to more careful and consistent grading during reinforcement learning). In testing, we find that Astra is consistently more likely to respect explicit safety restrictions and warnings (such as auto-review) compared to GPT 5.6 Sol.
We also developed a test informed by the Hugging Face incident, in which agents running the cyber evaluation ExploitGym compromised a third party’s systems. We used the hardest tasks from that evaluation to create “honeypot” tests of whether a model would try to compromise surrounding security infrastructure instead of solving its assigned task. GPT‑5.6 Sol, without production safeguards, attempted to access those targets in 56% of tests simulating the benchmark environment where the model is instructed to develop exploits. Astra made no such attempts in these tests. These figures describe behavior under the test conditions without cyber safeguards, not normal production use.
Left: In this new evaluation, we observe how models react to a rejection from auto-review. Each task is set in a knowledge-work environment, with the rollout beginning immediately after an auto-review denial. There are multiple ways then for the model to proceed: end the rollout, attempt to find a safer alternative, retry a similar command, or exploit a_poor configuration_ _to bypass auto-review entirely. Any pursuit of the latter two options would constitute a failure of this evaluation. Astra never attempted to circumvent auto-review.
Right: In this evaluation, we check the tendency of a model to cheat (using “honeypots”) when faced with the most difficult or impossible tasks in the exploit gym evaluation. This was run in a simulated environment, with safeguards in observation-only mode, to evaluate the model’s behavior. While GPT‑5.6 Sol indicated higher likelihood for cheating in these scenarios, Astra did not attempt to use those shortcuts, and it was still able to legitimately solve some tasks._
Training models to be aligned is one layer of safety. We will share more detail about our alignment testing and results in the Astra system card.
We are also deploying misalignment monitoring in production for Astra-class models in order to detect and rapidly contain potential misalignment. These safeguards resemble our monitoring for internal deployments and involve a system of classifiers which check the model’s reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity. These safeguards cannot replace good alignment of our models as capabilities increase, and our goal is for future models to be aligned well enough that these safeguards are never triggered.
What this will mean for users
OpenAI is committed to ensuring that the benefits of AI are broadly accessible. Given the significant increase in Astra’s cybersecurity capabilities, we are being especially careful to make this deployment safe and secure. Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity.
The system may occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped. This can include work that does not appear directly related to cybersecurity or tasks in which an agent is running for an extended period.
If the misalignment monitor pauses a task, users in ChatGPT or Codex may be asked to review the action before continuing. When using other surfaces like the API, the task will stop. We plan to keep calibrating these safeguards to reduce unnecessary interruptions and expand access to frontier capabilities through programs like Daybreak.
Looking forward
We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects. Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow.
That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient.
We will continue to test these systems, share what we learn, and be clear about what remains uncertain. The models that follow Astra will demand more of us. We will take the time and do the work needed to meet that responsibility.