将前沿智能融入你日常工作的更多方式。
本月早些时候,我们推出了 GPT‑6 Astra,这是世界上最智能、对齐程度最高的模型。尽管最严苛、最重要的项目仍然需要 Astra 的完整深度,但工作会以不同的规模、节奏和预算发生。
因此,我们通过 GPT‑6 Sol 和 GPT‑6 Luna 扩展了 GPT‑6 宇宙。GPT‑6 Astra 开启了新一代智能——这些模型通过在成本效率上推进前沿,帮助将这种智能的益处更广泛地分配。我们使用与 GPT‑6 Astra 类似的方法训练了 GPT‑6 Sol 和 Luna,将 Astra 在专业工作、事实性、编程、计算机使用和对齐方面取得最先进表现背后的进步,带到了更快、更实惠的模型中。
GPT‑6 系列模型在成本—智能曲线上全面领先,在每一层级都具备卓越能力,并配有能够高效大规模交付这些能力的基础设施。缓存和推理方面的改进让我们能够以更低的成本提供这些模型,我们通过将 Sol 和 Luna 的 API 价格相较其 GPT‑5.6 促销定价降低 50%,把这些节省直接传递给用户和客户。这些改进共同让先进 AI 在更多日常任务和大规模应用中变得切实可行。
GPT‑6 API 定价
模型 | 输入 | 输出 | 降价 |
GPT‑6 Sol | $4 → $2 | $20 → $10 | 便宜 50% |
GPT‑6 Luna | $0.20 → $0.10 | $1.20 → $0.50 | 便宜 50% |
价格按每 100 万 tokens 计。
GPT‑6 Astra 依然是我们全方位表现最好的模型。当你想要最佳结果和不妥协的体验时,选择它。
整个模型家族全面升级
GPT‑6 Sol 和 Luna 为你已经熟悉并使用的模型带来智能升级和成本效率,覆盖对完成复杂工作最有用的各项能力。
专业工作
GPT‑6 Sol 能够承担困难的工作任务,同时凭借更高的使用限额和更低的成本,给你留出更多迭代空间,相比价格相近的竞品模型提供更强的智能和更好的结果。
在 AutomationBench 这项跨应用业务流程测试中,GPT‑6 Sol 在 xhigh effort 下超越了 max effort 的 Claude Opus 5,而每任务成本仅为 Opus 5 的 9%。在 high effort 下,GPT‑6 Luna 相比其前代提升 5.4 个百分点,每任务成本降低 58%。
在 AutomationBench 1.0.6 中,AI 智能体使用 47 种工具,在销售、营销、运营、支持、财务和人力资源领域的端到端工作流上接受测试。Claude Fable 5.1 的数据点低估了其实际成本,因为它省略了 Opus 5 回退的成本,而这类回退在约 40% 的任务上出现。
GPT‑6 Sol 也以远更低的成本超越 Claude Fable 5.1,甚至胜过 low-effort 的 GPT‑6 Astra。
模型(及 effort) | 得分 | 每任务成本 |
|---|---|---|
GPT‑6 Sol(xhigh) | 33.2% | $0.27 |
GPT‑6 Astra(low) | 30.3% | 3.9x GPT‑6 Sol |
Claude Opus 5(max) | 26.9% | 11.1x GPT‑6 Sol |
Claude Fable 5.1 搭配 Opus 5 回退(max) | 31.4% | >8.9x GPT‑6 Sol (未报告回退成本) |
在 Agents’ Last Exam 上,该评测针对复杂专业工作流对智能体进行评估,GPT‑6 Sol 在最大努力程度下得分 56.4%,高于 Claude Opus 5 在该评测中的最高分,且每任务成本低 60%。
在 Agents’ Last Exam V1 中,AI 智能体在跨越 55 个子行业、涵盖在计算机上执行的大多数主要专业工作领域的长期、具有经济价值的任务上接受评估。
事实性
答案的有用性取决于事实是否正确,我们正在持续提升事实可靠性。在我们的内部事实性评测中,该评测基于去标识化的真实世界对话,其中用户标记了我们模型所犯的错误,GPT‑6 Sol 所犯的错误约为其前代的一半,以低得多的成本接近 Astra 级别的可靠性。GPT‑6 Luna 也有大幅提升;在更高的努力程度下,它以约百分之一的成本达到 GPT‑5.6 Sol 的水平。
这里我们在去标识化的 ChatGPT 对话上评估事实性,这些对话中用户曾标记过来自先前模型的事实错误。这些诱发错误的对话并不代表典型使用情况,在典型使用中事实错误更为罕见。得分未对长度进行控制;不过,我们的冗长度扫描显示答案长度几乎不产生影响。
编程
今年,编程智能体开始处理比以往任何时候都更复杂、范围更大、持续时间更长的任务。在 OpenAI,我们的内部使用量呈指数级增长。按 API 价格计算,每日 token 使用量中位数研究员已超过 $600,第 90 百分位的研究人员则超过 $7,000(研究加速:OpenAI 内部视角)。随着编程智能体承担更长、更高要求的任务,持续使用的成本变得更加重要。GPT‑6 Sol 和 Luna 将强大的编程性能与更低的 API 价格相结合,让开发者有更多空间进行迭代,也让团队有信心对 Codex 提出更宏大的要求。
在 FrontierCode 上——该基准评估编程智能体生成的改动是否已准备好合并进真实代码库——GPT‑6 Sol 相比 GPT‑5.6 Sol 有大幅提升,并且能够以低得多的成本匹配 Claude Fable 5.1 xhigh。
在 FrontierCode 1.1 Main 中,AI 智能体编写的代码不仅按正确性评分,还按“可合并性”评分:例如测试质量、范围把控、代码风格,以及对代码库规范的遵循程度。
在 DeepSWE v1.1 上,该基准测试在真实代码库中执行复杂软件工程任务的表现,GPT‑6 Sol 在最大努力下得分 68.8%,与 Claude Fable 5 在该评测中的最高分——xhigh 努力下的 69.9%——相差 1.1 个百分点,而每任务成本约低 80%。
GPT‑6 Luna 在最大努力档位下得分 66.6%,与 Claude Opus 5 和 Fable 5 在中等努力档位下的表现相当。在这些对比中,Luna 每项任务的成本比 Opus 5 低 93%,比 Fable 5 低 96%。
在 DeepSWE 1.1中,AI 智能体解决原创的、长时程软件工程任务。
计算机使用
尽管 GPT‑6 Astra 仍是全球最强的计算机使用模型,GPT‑6 Sol 和 Luna 相比前代提供了更具成本效益的性能。在 OSWorld 2.0 离线上,GPT‑6 Sol 在 xhigh 努力档位下取得了与 Claude Opus 5 在中等努力档位下相近的得分——60.5% 对 60.3%——而每项任务成本约低 80%。GPT‑6 Luna(max)能够以 GPT‑5.6 Sol(medium)十分之一的成本超越后者。
在 OSWorld 2.0中,AI 智能体尝试跨越日常与专业任务的长时程计算机使用工作流。我们报告的是 v2026.08.08 版本离线集上的部分奖励。
协作风格
我们还将 GPT‑6 Astra 改进后的沟通风格带到了 Sol 和 Luna 上,我们认为这在与技术和编程相关的对话中会尤为明显。可以期待更清晰的表达、更少的术语、更少古怪的措辞、更少低价值的细节,以及整体上略微更简短的回答,同时不损失实质内容。
当然可以。我可以在不引入 React 的情况下实现便当盒式的观感以及右上角的滑动页面切换器——这个交互足够小,在当前网站上就能保持快速流畅。我现在正围绕这套系统重新塑造全部四个页面。
完成了——打开线上网站看看。
现在全部四个页面都采用了色彩丰富的便当盒式布局,右上角还有一个带动画的 Home / Play / Place / Taste 滑块。并不需要 React;原生页面过渡让它保持轻量而流畅。
内置图像工具生成了这些食物插画。提示词:“一幅充满活力的水粉与水墨风格的编辑插画,赞美饺子和辣面,色彩明亮,有蒸汽、辛辣感,且没有文字或标志。”
我会把这些页面改造成便当盒网格布局,并把右上角的链接变成一个滑动的页面切换器。我会保留现有站点,并在改动其设置之前先确认这个交互是否需要 React。
更新后的站点已上线。现在每个页面都有活泼的便当盒布局,右上角的页面切换器会在各个区块之间滑动。我在桌面端和窄屏移动端上都做了检查,包括浏览器的后退导航。现有站点并不需要 React 来实现这个。
尽管风格是主观的,但我们在这里更偏好 GPT‑6 Sol 的回复。它不会那么快下结论,花更少时间重复对提问者来说可能显而易见的细节(例如网站有四个页面),使用更少含糊的措辞(例如“便当盒感”“围绕那个系统进行改造”),更坦率地说明自己检查了什么、没检查什么,也不会不必要地分享诸如其图像工具提示词之类的实现细节。
改进智能体和长对话的缓存
除了更低的 token 价格,我们还在帮助基于 GPT‑6 构建的开发者在其应用复用的上下文上节省更多。我们改进了 GPT‑6 的提示词缓存,默认提供更高的缓存命中率,帮助智能体复用更多上下文、更快响应,并享受缓存输入 token 读取 90% 的折扣。
开发者还有更多方式来衡量和优化其缓存性能:
监控与诊断。提示词缓存仪表盘会显示有多少输入被缓存,以及这一情况随时间的变化。诊断工具有助于解释错失的缓存机会以及需要修复的地方。
调整推理投入和工具可用性,而不破坏缓存。对于更难的任务,提高推理投入;对于更简单的后续任务,则降低它,并且随着你的智能体需求变化,启用或禁用工具。这两项控制现在都会保留更早的上下文以供缓存复用。
优化哪些前缀会被缓存。显式断点让开发者可以选择缓存提示词前缀的结束位置。这让开发者对缓存复用拥有更多控制权,并可以提升性能。
GitHub 报告称,在过去几个月里,这些改进使 OpenAI 模型数十亿次请求中需要全新处理的提示词 token 占比降低了 50% 以上,从而帮助 Copilot 更快响应。
持续改进对齐
GPT‑6 Sol 和 Luna 建立在 Astra 引入的对齐工作之上,Astra 是我们迄今为止对齐程度最高的模型。在我们的对齐评估中,Sol 和 Luna 都相较其 GPT‑5.6 对应版本有所改进,包括降低了就其编码工作作出误导性声明的比率。
以下评估刻意测试了具有挑战性的场景,并不衡量典型使用中的失败率。完整结果请参见系统卡。
在我们的内部编码欺骗评估中,AI 智能体被赋予刻意挑选以诱发不诚实行为的任务。在典型使用中,欺骗行为要罕见得多。欺骗率衡量的是检测到任何欺骗行为的回答所占的比例。推理努力程度设为最大。
可用性
GPT‑6 Sol 和 GPT‑6 Luna 从今天起在 ChatGPT Work 和 Codex 中面向所有 Plus、Pro、Business、Enterprise 和 Edu 用户开放。Free 和 Go 用户可以在桌面应用中访问 GPT‑6 Luna。这些模型尚未在 Chat 中提供。在 OpenAI API 中,它们以gpt-6-sol和gpt-6-luna的形式提供。
为了保持对所有人的服务稳定,我们计划在一天内逐步在 ChatGPT 中推出这些模型。如果你在 ChatGPT Work 或 Codex 中没有看到新模型,请稍后再试。
GPT 的评估是在我们的研究环境中或通过我们的 API 进行的,由于系统提示词、可用工具等方面的差异,其输出可能与生产环境中的 ChatGPT 略有不同。竞品模型的评估取自公开报告。在无法获得 Claude Fable 5.1 的分数时,报告的是 Claude Fable 5 的分数。
More ways to bring frontier intelligence into the work you do every day.
Earlier this month, we introduced GPT‑6 Astra, the most intelligent and aligned model in the world. While the most demanding and important projects still call for Astra’s full depth, work happens at different scales, rhythms, and budgets.
That’s why we’re expanding the GPT‑6 universe with GPT‑6 Sol and GPT‑6 Luna. GPT‑6 Astra introduced a new generation of intelligence—these models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency. We trained GPT‑6 Sol and Luna with similar methods as GPT‑6 Astra, bringing the advances behind Astra’s state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models.
The GPT‑6 models lead across the cost–intelligence curve, combining exceptional capabilities at every tier with infrastructure that delivers them efficiently at scale. Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers by reducing API prices for Sol and Luna by 50% compared with their GPT‑5.6 promotional pricing. Together, these improvements make advanced AI practical for more everyday tasks and applications at scale.
GPT‑6 API pricing
Model | Input | Output | Price reduction |
GPT‑6 Sol | $4 → $2 | $20 → $10 | 50% cheaper |
GPT‑6 Luna | $0.20 → $0.10 | $1.20 → $0.50 | 50% cheaper |
Prices are per 1 million tokens.
GPT‑6 Astra continues to be our best model across the board. Choose it when you want the best results and an uncompromising experience.
A step up across the model family
GPT‑6 Sol and Luna bring intelligence upgrades and cost efficiency to the models you already know and use across capabilities most useful for getting complex work done.
Professional work
GPT‑6 Sol can take on difficult work tasks while giving you more room to iterate with higher usage limits and lower cost, offering more intelligence and better results versus similarly priced competitor models.
On AutomationBench, a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task. At high effort, GPT‑6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.
In AutomationBench 1.0.6, AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. The datapoint for Claude Fable 5.1 understates its actual cost, as it omits the cost of the Opus 5 fallbacks, which occurred on ~40% of tasks.
GPT‑6 Sol also exceeds Claude Fable 5.1 at far lower cost, and even bests low-effort GPT‑6 Astra.
Model (and effort) | Score | Cost per task |
|---|---|---|
GPT‑6 Sol (xhigh) | 33.2% | $0.27 |
GPT‑6 Astra (low) | 30.3% | 3.9x GPT‑6 Sol |
Claude Opus 5 (max) | 26.9% | 11.1x GPT‑6 Sol |
Claude Fable 5.1 w/ Opus 5 Fallback (max) | 31.4% | >8.9x GPT‑6 Sol (fallback cost not reported) |
On Agents’ Last Exam, which evaluates agents on complex professional workflows, GPT‑6 Sol at max effort scores 56.4%, above Claude Opus 5’s highest score in the evaluation at 60% lower cost per task.
In Agents’ Last Exam V1, AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries, covering most major fields of professional work performed on a computer.
Factuality
The usefulness of an answer depends on getting the facts right, and we’re continuing to make progress on factual reliability. On our internal factuality evaluation, which is based on de-identified real-world conversations where users flagged mistakes by our models, GPT‑6 Sol makes about half as many mistakes as its predecessor, approaching Astra-level reliability at much lower cost. GPT‑6 Luna also improves substantially; at higher effort levels it matches GPT‑5.6 Sol at about a hundredth its cost.
Here we evaluate factuality on de-identified ChatGPT conversations where users had flagged a factual error from a prior model. These error-inducing conversations are not representative of typical usage, where factual errors are more rare. Scores are not controlled for length; however, our verbosity sweeps showed almost no dependence on answer length.
Coding
This year, coding agents have begun tackling tasks with more complexity, scope, and duration than ever before. At OpenAI, our internal usage has grown exponentially. Valued at API prices, daily token usage has exceeded $600 for the median researcher and $7,000 for researchers at the 90th percentile (Research acceleration: The view inside OpenAI). As coding agents take on longer and more demanding tasks, the cost of sustained use matters more. GPT‑6 Sol and Luna combine strong coding performance with lower API prices, giving developers more room to iterate and teams the confidence to be more ambitious about what they ask Codex to take on.
On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.
In FrontierCode 1.1 Main, AI agents write code that’s graded not only on correctness but also “mergeability”: e.g., test quality, scope discipline, code style, and adherence to codebase standards.
On DeepSWE v1.1, which tests performance on complex software-engineering tasks in real codebases, GPT‑6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.
GPT‑6 Luna at max effort scores 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort. In these comparisons, Luna costs 93% less per task than Opus 5 and 96% less than Fable 5.
In DeepSWE 1.1, AI agents solve original, long-horizon software engineering tasks.
Computer use
While GPT‑6 Astra remains the world’s best model for computer use, GPT‑6 Sol and Luna offer more cost-efficient performance than their predecessors. On OSWorld 2.0 offline, GPT‑6 Sol at xhigh effort achieves a similar score to Claude Opus 5 at medium effort—60.5% versus 60.3%—at approximately 80% lower cost per task. GPT‑6 Luna (max) is able to exceed GPT‑5.6 Sol (medium) at one tenth of its cost.
In OSWorld 2.0, AI agents attempt long-horizon computer-use workflows spanning everyday and professional tasks. We report the partial reward on the offline set from the v2026.08.08 release.
Collaboration style
We’ve also brought GPT‑6 Astra’s improved communication style to Sol and Luna, which we think will be especially noticeable in technical and coding conversations. Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance.
Absolutely. I can get the bento feel and the top-right sliding page switcher without pulling in React—the interaction is small enough to stay fast and smooth in the current site. I’m reshaping all four pages around that system now.
Done — open the live site.
It now has a colorful bento-box layout across all four pages and an animated top-right Home / Play / Place / Taste slider. React wasn’t necessary; native page transitions keep it lightweight and smooth.
The built-in image tool created the food artwork. Prompt: “An energetic gouache-and-ink editorial illustration celebrating dumplings and spicy noodles, with bright colors, steam, spice, and no text or logos.”
I’ll reshape the pages into a bento grid and turn the top-right links into a sliding page switcher. I’ll keep the existing site and check whether that interaction needs React before changing its setup.
The updated site is live. Every page now has a playful bento layout, and the top-right page switcher slides between sections. I checked it on desktop and narrow mobile screens, including browser back navigation. The existing site didn’t need React for this.
Although style is subjective, we prefer GPT‑6 Sol’s reply here. It doesn’t jump to conclusions as quickly, spends less time reiterating details that might be obvious to the asker (e.g., that the website has four pages), uses less vague language (e.g., “bento feel”, “reshaping… around that system”), is more forthcoming with what it did and didn’t check, and doesn’t unnecessarily share implementation details like its image tool prompt.
Improving caching for agents and long conversations
Alongside lower token prices, we’re helping developers building on GPT‑6 save more on the context their applications reuse. We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.
Developers also have more ways to measure and optimize their caching performance:
Monitor and diagnose. The Prompt Caching Dashboard shows how much input is cached and how that changes over time. The diagnostics tool helps explain missed opportunities for caching and what to fix.
Adjust reasoning effort and tool availability without breaking cache. Increasereasoning effort for harder tasks or lower it for simpler follow-ups, andenable or disable tools as your agent’s needs change. Both controls now preserve earlier context for cache reuse.
Optimize which prefixes get cached. Explicit breakpoints let developers choose where cached prompt prefixes end. This gives developers more control over cache reuse and can improve performance.
GitHub reports that, over the past several months, these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models, helping Copilot respond faster.
Continuing to improve alignment
GPT‑6 Sol and Luna build on the alignment work introduced with Astra, our most aligned model to date. In our alignment evaluations, both Sol and Luna show improvements over their GPT‑5.6 counterparts, including lower rates of misleading claims about their coding work.
The evaluations below deliberately test challenging situations and do not measure failure rates in typical use. See the system card for the full results.
In our internal coding deception evaluation, AI agents are given tasks deliberately selected to elicit dishonesty. In typical usage, deception is much rarer. Deception rate measures the fraction of answers with any detected deception. Effort was set to maximum.
Availability
GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access GPT‑6 Luna in the desktop app. These models are not yet available in Chat. In the OpenAI API, they are available as gpt-6-sol and gpt-6-luna.
To keep service stable for everyone, we plan to roll out these models in ChatGPT gradually throughout the day. If you don’t see the new models in ChatGPT Work or Codex, please try again later.
Evaluations of GPT were performed in our research environment or via our API, which may provide slightly different output from production ChatGPT due to differences in the system prompts, tools available, etc. Evaluations of competitor models were taken from publicly available reports. Scores for Claude Fable 5 were reported when scores for Claude Fable 5.1 were unavailable.