我们最强大的模型,专为企业所需完成的各种工作而打造。

上周,我们推出了 GPT‑6 Astra——全球最智能、对齐程度最高的模型,现已上线 ChatGPT Work、Codex 和 API。Astra 在计算机使用、浏览、专业工作、软件工程、网络安全和科学领域均达到 SOTA 水平,让团队能够以无与伦比的速度、准确性和判断力应对最具挑战性的专业工作。
全球处理复杂工作的最佳模型
大多数 AI 系统都要求企业先准备数据、重新设计工作流程并构建自定义集成,然后才能产生价值。Astra 改变了这一切。在 ChatGPT Work 和 Codex 中,它能够编写代码,并像人一样操作人们日常使用的各类应用程序——即使这些应用程序没有 API 也能做到。这意味着企业从第一天起就能让 AI 在现有工作流程中发挥作用,无需大量前期准备或工程工作。
Excel 竞赛:****GPT‑6 Astra 可通过计算机使用方式完成 Financial Modeling World Cup 挑战,速度约为获胜人类选手的四倍——帮助分析师减少搭建模型的时间,把更多时间用于解读结果和做出决策。来自**2023 年 Microsoft Excel 世界锦标赛** .
在正式上线后的最初几天里,我们已经看到客户开始让 Astra 投入实际工作,从优化 GPU、发现财务报表中的差异,到制作更贴合品牌形象的演示文稿,不一而足。
“我们将在 Devin 正式发布当天就把 GPT‑6 Astra 集成到其 harness 中,它在我们的内部测试基准上展现出了业界领先的性能。它出色的计算机使用能力、写作能力和代码库理解能力让测试体验即刻得到提升:视频明显更容易跟读了,报告也更清晰、更简洁。”
—— Silas Alberti,Cognition 研究高级副总裁
“GPT‑6 Astra 使用我们的 Genie harness,在 OfficeQA Pro 和 Pro V2 基准上取得了新的业界最佳成绩。与 GPT‑5.6 Sol 相比,它的单任务成本也显著更低。在我们的全部基准测试中,它在面向企业的数据推理和文档理解方面都展现出明显的跃升。”
—— Ivan Zhou,Databricks 资深研究工程师兼技术负责人
“Astra 在我们的评测中创下新高:它生成的演示文稿是我们测试过的最好的,对简报要求的遵循度比次优模型高出 17%,并且引用来源指向正确文档的频率高出 19%。对于需要处理大量密集金融材料的分析师来说,这意味着你生成的演示文稿和答案可以直接交给客户,并且能够逐条为其辩护。”
—— George Sivulka,Hebbia 创始人兼首席执行官
“Astra 是我们测试过的最强模型之一,在复杂企业工作流中展现出领先性能。最突出的是它的判断力——它更擅长拒绝断言文档并不支持的结论,在整个评估中,它做出自信但错误断言的可能性降低了超过 10%。”
—— Yashodha Bhavnani,Box 公司 AI 产品副总裁
“Astra 能理解你的构想,并知道如何运用 Figma 去实现它,在复杂设计中推进工作的同时,让你始终掌握创意方向的主导权。”
—— Loredana Crisan,Figma 首席设计官
“Astra 在智能水平和写作质量上实现了明显跃升,具备更强的多智能体协调能力,也能更好地把握请求背后未言明的意图。这种组合在法律工作中至关重要,因为理解用户想要达成的目标,并以精准、细腻的方式将其表达出来,是不可或缺的能力。”
—— Omar Bari,Thomson Reuters Labs 应用研究副总裁
- Cognition
- Databricks
- Hebbia
- Box
- Figma
- Thomson Reuters Labs
在 OpenAI,Astra 在正式发布前数周便已在内部铺开,因此我们亲眼见证了计算机使用等前沿能力如何改变我们的工作方式。我们的开发团队和营销团队利用 Astra 和 Codex,将三个小时的多机位素材制作成了 GPT‑6 Astra 开发者初体验视频 ,该视频上线仅 4 天就已获得超过 55 万次观看。我们的工程团队利用 Astra 发现并解决了一个内存分配瓶颈,该瓶颈曾导致测试环境中 Codex 会话响应缓慢。通过切换分配器,他们将单轮延迟降低了 25 倍,同时峰值内存占用仅增加约 30%。
Astra 在遵循公司语气、模板和设计规范方面也表现更佳,因此首次生成的结果就更接近团队可直接投入使用的成品。
参考文件
GPT‑6 Astra 输出
GPT‑6 Astra 仅使用 OpenAI 演示文稿模板中的几张幻灯片,就制作出了一个关于虚构模型 GPT‑Gaia 的幻灯片演示,并在整个过程中保持了正确的语气和版式。这意味着你可以获得符合企业标准、格式正确的演示文稿。
每一美元都能完成更多有价值的工作
Astra 延续了我们提供极致高效模型的承诺,让客户每一美元都能获得更多有效产出。它经过训练,能用更少的 token 和更少的重试次数完成任务,这意味着更少的返工和更低的单任务成本。借助 Astra,OpenAI 在专业工作和编程评测(包括 Terminal Bench 4.0 和 Artificial Analysis Intelligence Index)上占据了成本效率前沿的大部分位置。定价为每百万输入 token 10 美元起,每百万输出 token 50 美元起。
Terminal-Bench 4.0 测试智能体在复杂终端任务上的表现,包括软件工程、系统配置和数据分析。GPT‑6 Astra 以 57.9% 的成绩创下新高,相比之下 GPT‑5.6 Sol2为 37.3%,Claude Fable 5.1 为 55.8%,而 Astra 的每任务估算 API 成本分别低约 9% 和 63%。
“Astra 在 DeepSWE v1.1 上以 74% 的成绩创下新纪录。它用比前沿模型更少的步骤和更高的 token 效率做到了这一点,尤其是在复杂、长周期的任务上。毫无疑问,这个模型将对高质量的真实世界软件工程产生显著影响。”
— Serena Ge,Datacurve 联合创始人兼 CEO
“在我们的主动式智能体工作流中,我们看到 Astra 将耗时 5 小时以上的端到端工作流的通过率提升了 20%,同时减少了完成任务所需的推理调用次数。该模型更强的决策能力让我们能够移除脚手架并加速性能表现。”
— Mitch Troyanovsky,Basis 联合创始人
“与我们的基线相比,Astra 捕获的 bug 数量多出约 20%。在那些需要跨文件深度推理才能发现细微问题的 pull request 上,它的捕获率提高了一倍以上。在代码审查中,它能将变更的意图与其后果联系起来:它跨文件推理,捕获基线模型遗漏的接口契约漂移和授权漏洞,并用具体的验证步骤支撑其发现。”
— David Loker,CodeRabbit AI 副总裁
“过去几个月,观察 AI 数学研究能力的跃升令人着迷。我们正处于这样一个时期:每一款新模型都在真正推动可能性的边界。Astra 是又一个显著的进步,我预计未来两年内进展速度将大幅加快。”
— Alex Gerko,XTX Markets CEO
- Datacurve
- Basis
- CodeRabbit
- XTX Markets
为高影响工作提供更多安全性与可控性
让 AI 访问业务系统,需要对其行为方式有充分的信心。有鉴于此,Astra 是我们迄今对齐程度最高的模型,对人类意图和授权的遵循更加严格。
在训练过程中,我们在 内部计算机使用安全基准 上对 Astra 进行了测试,该基准针对最困难的业务场景(如泄露机密信息、过度广泛地共享仪表板或删除数据)来评估模型。在这项评估中,Astra 产生意外结果的频率比 GPT‑5.6 Sol 低 89%,比 Claude Fable 5.1 低 74.7%。额外的确认机制和自动审查进一步提升了 GPT‑6 Astra 和 GPT‑5.6 Sol 的性能。
组织还可以决定 Astra 的部署范围。新的企业管理控制功能允许他们限制对已批准网站和桌面应用程序的访问、管理上传和下载,并控制浏览历史。ChatGPT Work 和 Codex 也包含安全防护措施,例如确认策略(可在执行重大操作前要求审批)以及对潜在不安全或未经授权的工具调用进行自动审查。这些控制功能使团队能够从受限配置起步,并随时间推移逐步扩大访问范围。
为进一步扩大访问范围,除 Astra 外,我们还在 ChatGPT Desktop 中推出了新的企业插件。由最新的浏览器使用能力驱动,来自 Oracle Analytics、Power BI(Microsoft Fabric 服务)、Navan 和 Avalara 的插件让访问熟悉的企业应用程序变得更加容易。
Astra也是首个达到关键网络安全能力阈值(依据我们的预备框架)的模型。随着能力的提升,我们加强了防护措施,以应对滥用行为以及模型未经授权采取行动的风险,包括训练 Astra 尊重安全与安保边界、提升其抵御规避安全机制尝试的能力,以及部署旨在阻止有害响应的自动化检查。
零数据保留(Zero Data Retention)适用于符合条件的 API 客户,在受支持的端点上可用,且需经审批。
立即开始使用 Astra
在 ChatGPT Work 或 Codex 中试用 GPT‑6 Astra,或通过 API 将其集成到你自己的产品与工作流中。
企业管理员可在适用的费率卡与协议下启用 Astra。上线时企业访问默认处于关闭状态。
Our most capable model, built for all the work businesses need to get done.

Last week we introduced GPT‑6 Astra, the world’s most intelligent and aligned model, now available in ChatGPT Work, Codex, and the API. Astra is state-of-the-art on computer use, browsing, professional work, software engineering, cybersecurity, and science, so teams can take on the most demanding professional work with unmatched speed, accuracy, and judgment.
The world’s best model for complex work
Most AI systems require businesses to prepare their data, redesign workflows, and build custom integrations before they can deliver value. Astra changes that. In ChatGPT Work and Codex, it can write code and work through the same applications people use every day—even when those applications don’t have an API. That means businesses can put AI to work within their existing workflows from day one, without extensive preparation or engineering work.
Excel competition:****GPT‑6 Astra can complete Financial Modeling World Cup challenges using computer use aboutfour times as fastas the winning human competitor—helping analysts spend less time building models and more time interpreting results and making decisions. From the**2023 Microsoft Excel World Championship** .
Within the first few days of rollout, we’re already seeing customers put Astra to work, from optimizing GPUs to spotting discrepancies in financial statements to producing more on-brand decks.
“We’re integrating GPT‑6 Astra into Devin’s harness on launch day, where it delivers state-of-the-art performance on our internal testing benchmark. Its excellent computer use, writing, and codebase understanding improved testing right out of the box: videos are noticeably easier to follow, and reports are clearer and more concise”
— Silas Alberti, SVP Research, Cognition
“GPT‑6 Astra claims the new state of the art on our OfficeQA Pro & Pro V2 benchmarks, using our Genie harness. It also offers significantly better cost per task than GPT‑5.6 Sol. Throughout our benchmarks, it shows a clear step up on data reasoning and document understanding for enterprises.”
— Ivan Zhou, Staff Research Engineer & Tech Lead Manager, Databricks
“Astra set a new high in our evals: it produced the best decks we've tested and followed the brief 17% more faithfully than the next-best model, while sourcing its claims to the right document 19% more often. For analysts working through dense financial materials, that means decks and answers you can hand to a client and defend line by line.”
— George Sivulka, Founder & CEO, Hebbia
“Astra is one of the strongest models we’ve tested, delivering leading performance across complex enterprise workflows. What stood out most was its judgement — it was better at declining to assert conclusions the documents didn't support, and across the evaluation it was >10% less likely to make confidently incorrect assertions.”
— Yashodha Bhavnani, VP of AI Products, Box
“Astra gets your vision and knows how to use Figma to achieve it, working through complex designs while you stay in control of the creative direction.”
— Loredana Crisan, Chief Design Officer, Figma
“Astra delivers a clear jump in intelligence and writing quality, with stronger multi-agent coordination and a better grasp of the quiet intent behind a request. That combination matters in legal work, where understanding what the user is trying to accomplish and expressing it with precision and nuance are essential.”
— Omar Bari, VP Applied Research, Thomson Reuters Labs
- Cognition
- Databricks
- Hebbia
- Box
- Figma
- Thomson Reuters Labs
At OpenAI, Astra was rolled out internally weeks before launch, so we saw first hand how bleeding-edge capabilities like computer use could change the way we work. Our developer and marketing teams used Astra and Codex to turn three hours of multicamera footage into our GPT‑6 Astra Developer First Impressions video which has already garnered over 550k views in just 4 days. Our engineering team used Astra to uncover and resolve a memory-allocation bottleneck that was causing slow Codex sessions in a test environment. By switching allocators, they were able to produce 25× lower turn latency with roughly 30% higher peak memory use.
Astra is also better at following a company’s voice, templates, and design standards, so the first result is closer to something a team can put to use.
Reference file
GPT‑6 Astra output
GPT‑6 Astra creates a slideshow about GPT‑Gaia, a fictional model, using just a few slides from OpenAI’s presentation template, capturing the correct tone and layout throughout. This means you can expect slide decks that are correctly formatted for your business standards.
More useful work for every dollar
Astra continues our commitment to providing extremely efficient models that deliver more useful work per dollar to our customers. It's been trained to complete tasks in fewer tokens with fewer retries, which means less rework and lower cost per task. With Astra, OpenAI occupies the majority of the cost-efficiency frontier on professional work and coding evaluations, including Terminal Bench 4.0 and Artificial Analysis Intelligence Index. Pricing starts at $10 per million input tokens and $50 per million output tokens.
Terminal-Bench 4.0 tests agents on complex terminal-based tasks, including software engineering, system configuration, and data analysis. GPT‑6 Astra reaches a new high at 57.9%, compared with 37.3% for GPT‑5.6 Sol2and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task, respectively.
“Astra sets a new record on DeepSWE v1.1 at 74%. It did so with fewer steps and greater token efficiency than has ever been achieved by frontier models, especially on complex, long horizon tasks. Certainly, this model will have a noticeable impact on high quality, real-world software engineering.”
— Serena Ge, Co-Founder & CEO, Datacurve
“For our proactive agent workflows, we've seen Astra improve the pass rate in end-to-end workflows that take 5+ hours by 20%, while reducing the number of inference calls needed to complete the work. The model's improved decision making allows us to remove scaffolding and accelerate performance.”
— Mitch Troyanovsky, Co-founder, Basis
“Compared with our baseline, Astra caught ~20% more bugs. On pull requests that require extensive cross-file reasoning to detect subtle issues, it more than doubled the catch rate. In code review, it connects a change's intent to its consequences: it reasons across files to catch interface-contract drift and authorization bugs the baseline missed, and it backs findings with concrete verification steps.”
— David Loker, VP of AI, CodeRabbit
“It's been fascinating to observe the jumps in AI mathematical research capabilities over the last few months. We're in a period where each new model really pushes the frontier of what is possible. Astra is another notable step forward and I expect the pace of progress to massively accelerate over the next two years.”
— Alex Gerko, CEO, XTX Markets
- Datacurve
- Basis
- CodeRabbit
- XTX Markets
More safety and control for consequential work
Giving AI access to business systems requires confidence in how it will act. With this in mind, Astra is our most aligned model yet, with stronger adherence to human intent and authorization.
During training, we tested Astra on our internal computer use safety benchmark which tests models against the hardest business scenarios such as exposing confidential information, sharing a dashboard too broadly, or deleting data. In this evaluation, Astra produced unintended outcomes 89% less often than GPT‑5.6 Sol and 74.7% less often than Claude Fable 5.1. Additional confirmation and automated review further improved performance for GPT‑6 Astra and GPT‑5.6 Sol.
Organizations can also decide how broadly to deploy Astra. New enterprise admin controls let them restrict access to approved websites and desktop applications, manage uploads and downloads, and control browsing history. ChatGPT Work and Codex also include safeguards such as confirmation policies, which can require approval before consequential actions, and automated review of potentially unsafe or unauthorized tool calls. These controls allow teams to start with a limited configuration and expand access over time.
To further access, alongside Astra, we’re also launching new enterprise plugins in ChatGPT Desktop. Powered by the latest browser use capabilities, plugins from Oracle Analytics, Power BI (a Microsoft Fabric service), Navan, and Avalara make it easier to access familiar enterprise applications.
Astra is also the first model to reach the Critical cybersecurity capability threshold under our Preparedness Framework. With that increased capability, we’ve strengthened protections against both misuse and the model taking unauthorized actions including training Astra to respect safety and security boundaries, improving its resistance to attempts to bypass safeguards, and deploying automated checks designed to block harmful responses.
Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval.
Start using Astra today
Try GPT‑6 Astra in ChatGPT Work or Codex , or build it into your own products and workflows through the API .
Enterprise administrators can enable Astra under their applicable rate card and agreement. Enterprise access is off by default at launch.