我们正式推出 Claude Opus 5.5,它是我们全新 Claude 5.5 家族中的首个模型。在大多数工作上,它的表现达到 Claude Fable 5.1 的水平,而运行成本比 Opus 5 低 40%。
Claude Opus 5.5 是我们自呼吁 放缓前沿步伐 以来的首个发布版本。它在发布前由外部评估方进行了测试,包括 Frontier Design 和 METR。在我们运行的自动化行为审计——也就是我们所做的最全面的对齐测试——中,Opus 5.5 是我们迄今测试过表现最强的模型。它还配备了我们已经为最强能力模型开发的安全防护措施。
以下是你可以从 Opus 5.5 中期待的一些改进:
性能。 Opus 5.5 相较 Opus 5 是一次重大跃升。它是新的领先模型,早期测试者在其最复杂的工作上看到了性能的大幅提升。一位测试者在不到一天内完成了一项 68 万行的代码迁移——这项工作原本需要一个工程团队花费数周。它擅长发现并修复软件中的低效问题:当我们要求它缩短一个 Web 应用每个页面的加载时间时,Opus 5.5 在 40 次中成功了 39 次,而 Opus 5 只做出了较小的改进,而且还改变了应用的行为。另一位测试者让多个 Claude 模型根据单个提示词构建一款游戏;Opus 5.5 凭借其图形效果和精细度得分高于其他任何模型。
安全性。在我们迄今规模最大的自动化行为审计中——即那套在数千个模拟场景中对 Claude 进行测试的对齐套件——Opus 5.5 取得了所有模型中的最佳成绩。与近期模型相比,它采取难以逆转的行动或越出既定边界的可能性要低得多,而且比 Opus 5 更能抵御提示词注入。我们还扩大了对齐测试的范围,覆盖了更长的任务、不可能完成的任务,以及以真实事件为蓝本的场景,尽管它仍有局限。我们评估的完整细节可在 Opus 5.5 系统卡中查阅。
由于 Opus 5.5 在生物学和网络安全方面与 Claude Mythos 5.1 相当,我们为其部署配备了与 Claude Fable 5.1 类似的防护措施。经过审核的机构今天即可申请我们的 生命科学验证计划,将 Opus 5.5 用于生物学研究。未来几周内,我们还将扩大 网络安全验证计划的准入范围,经过验证的网络安全从业者将能够将 Opus 5.5 用于其工作。
成本与速度。Opus 5.5 在服务时所需的算力少于 Opus 5,其定价也反映了这一点。我们的测试显示,在默认设置下,它在典型工作负载上的成本将比 Opus 5 低 40%。输入和输出 token 分别为每百万 $4 和 $20,比 Opus 5 低 20%。缓存读取(占智能体和编程工作成本的大部分)为每百万 token $0.20,比 Opus 5 低 60%。Opus 5.5 生成输出的速度也比 Opus 5 快 30% 以上。
除了降价之外,我们还将提高 Pro、Max 和 Team 套餐的五小时使用限额。我们还会为订阅用户提供一次速率限制重置,你现在可以将其保存下来,并在你选择的任何时候使用。
沟通表达。 Opus 5.5 比之前的模型沟通得更自然。早期测试者发现它的写作更清晰、更易于理解,这解决了我们听到的关于 Opus 5 的一些常见反馈。它把最重要的信息放在前面,其风格使它成为长时间协作中更好的工作伙伴。正如一位早期测试者所说,“它写作的方式就像我一样。”在我们自己的使用中,这让 Opus 5.5 的工作成果更易于跟进和核查——这既是安全方面的好处,也是实用方面的好处。
Claude Sonnet 5.5 和 Claude Haiku 5.5 将在未来几周内推出,它们在性能、效率和安全性方面将带来许多相同的改进。
性能与成本效益
在我们的基准测试中,Claude Opus 5.5 在智能体编程、计算机使用和知识工作方面处于领先。话虽如此,在这些能力水平上,我们发现基准测试的差距已不再是衡量现实世界差异的可靠指南。在我们自己的使用中,Opus 5.5 与 Claude Fable 5.1 之间的差距比这些分数所显示的要小。
| Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol | |
|---|---|---|---|---|---|
| 智能体编程Terminal-Bench 4.0¹ | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| 智能体编程FrontierCode v1.1(主) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| 智能体编程CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| 知识工作GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| 业务流程自动化Bench² | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| 多学科推理Humanity's Last Exam | 67.7%使用工具 | 65.6%使用工具 | 63.6%使用工具 | 57.2%使用工具 | — |
| 智能体科学研究Terminal-Bench-Science 0.1³ | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| 计算机使用OSWorld 2.0 | 81.8%部分 | 80.7%部分 | 74.0%部分 | — | — |
| 可视化图表识别Chartography | 89.0%使用工具 | 88.4%使用工具 | 83.4%使用工具 | — | — |
除非另有说明,所有 Claude Opus 5.5 的结果均在最大努力下使用自适应思考。Terminal-Bench 4.0 的结果报告为 Claude Opus 5.5 在 xhigh 努力下以及 GPT-6 Astra 在 high 努力下的表现,数据由 OpenAI 报告;这些代表每个模型的最高得分。Claude Opus 5.5 在启用其生产防护措施的情况下进行了评估。当这些防护措施介入时,网络安全任务由 Claude Opus 4.8 完成,生物学和前沿 LLM 开发任务由 Claude Opus 5 完成。这很可能降低了 Claude Opus 5.5 在这些基准测试上的表现。
1 Terminal-Bench 4.0:Claude Opus 5.5 的标准误差为 ±2.6 分,其他 Claude 模型为 ±1.6–2 分。公开排行榜(每任务 5 次试验,Claude Code 测试框架)报告 Claude Opus 5 为 51.8%;我们的设置复现结果为 52.3%,在噪声范围内。GPT-6 Astra 和 GPT-5.6 Sol 的数据由 OpenAI 报告。
2 AutomationBench:AutomationBench 的结果由 Zapier 运行并报告。这些运行未使用回退模型,因此防护措施的介入被视为失败——这导致得分低于 Claude Opus 5.5 在实际中会取得的成绩。Claude Opus 5.5 的结果来自 Zapier 在早期访问期间自行进行的评估。Opus 5、GPT-5.6 Sol 和 GPT-6 Astra 的结果来自 Zapier 的公开排行榜。
3 Terminal-Bench-Science 0.1:每个模型的标准误差为 ±3.5–5 分。公开排行榜(每任务 3 次试验,Claude Code 测试框架)报告 Claude Opus 5 为 30.0%;我们的设置复现结果为 29.0%,在噪声范围内。GPT-6 Astra 的数据由 OpenAI 报告。
Opus 5.5 的优势非常明显的地方在于效率。它的每 token 成本低于 Opus 5,且每个任务使用的 token 更少,最终成本下降了 40%。
定价
| 每 1M tokens 价格 | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| 缓存读取 | $0.20 | $0.50 |
| 输入 tokens | $4 | $5 |
| 输出 tokens | $20 | $25 |
| 缓存写入 | $5 | $6.25 |
Opus 5.5 的快速模式也可在 Claude Code 和 Claude Platform 中使用,速度最高可达 2.5 倍。其价格为每百万输入 token 8 美元,每百万输出 token 40 美元。
编程
Opus 5.5 尤其擅长处理漫长而庞杂的任务,比如代码库范围的迁移和审计。一位早期测试者用它审计并修复了一个 20 万行的代码库,用时不到三小时,而 Opus 5 则耗时超过 20 小时,且使用的 token 数量是前者的 2.5 倍。
在一项内部测试中,我们让 Opus 5.5 和 Fable 5.1 将 HAProxy(一款广泛使用的、用于在多台服务器之间平衡网络流量负载的软件)从 C 语言翻译成 Rust。两个重写版本都通过了 HAProxy 自身几乎所有的回归测试,但 Opus 5.5 在 9.5 小时内完成,而 Fable 5.1 用了 12 小时,且成本低了 51%。
Opus 5.5 以极低的成本在智能体编程方面取得了前沿成果。在 FrontierCode 上以其默认投入程度运行时,它击败了 GPT-6 Astra,而每个任务的成本仅约为后者的 20%。在 Terminal Bench 4.0 上,它以约 40% 的成本与 Astra 持平,而在 CursorBench 上,它以约三分之一的成本领先 GPT-5.6 Sol 11 分。
Terminal-Bench 4.0 衡量的是模型在命令行界面内完成复杂、多步骤专业任务的能力。Opus 5.5 在默认投入程度下击败了以最大投入程度运行的 Opus 5,而成本仅约为后者的五分之一。它以约 40% 的成本与 GPT-6 Astra 持平。
我们的早期测试者报告了类似的效率和智能提升:
引用
“开发者想要的是能够承担真实软件工作并把它完成的智能体。在我们横跨 GitHub Copilot CLI 和 VS Code 的测试中,Claude Opus 5.5 使用的 token 和步骤数在我们所测量的模型中属于最少的之一。在 VS Code 中,它解决的终端任务比 Opus 5 更多,而步骤数不到后者的一半。它不只是让单个任务更高效,更是让开发者更大的项目变得更可实现。”
公司
GitHub
作者
Mario Rodriguez,首席产品官
最安全的编码智能体
在其系统内使用智能体的企业需要知道这些智能体正按预期运行,尤其是当它们自主运行许多小时的时候。Opus 5.5 有一个分类器,会在每个动作运行前对其进行筛查;有一个开源沙箱,安全团队可以对其进行审计;还有代码审查,能在漏洞合并之前将其捕获。
该模型本身也具备更强的防御能力。在提示词注入攻击方面,它在我们测试的每一种场景中都追平或超越了 Opus 5,包括编码、工具使用、计算机使用和网页浏览。在 AI 安全公司 Gray Swan 运行的一项基准测试中,Opus 5.5 与 Fable 5.1 并列,创下所有受测模型中最低的提示词注入成功率。
知识工作
Opus 5.5 是一位可靠且娴熟的研究员。在一项内部测试中,我们要求 Opus 5.5、Fable 5.1 和 Opus 5 仅凭它们在一个财报难以找到的网络副本上所能找到的信息,撰写一份关于某公司季度业绩的报告。一个自动评分器对照来源核查了每一个数字和引文。
在不同的 effort 设置下,Opus 5.5 的报告有 18 份中的 16 份达到了我们的质量门槛,而任何捏造的数字或引文都会导致不合格。Fable 5.1 和 Opus 5 在任何一次尝试中都未能达到该门槛。
它在财务分析和商业工作方面同样表现出色。投资公司 Walleye Capital 作为早期测试方报告称,Opus 5.5 在其最低设置下就基本解决了他们的评估套件;在更高设置下,它的表现更佳,甚至发现了他们评估指令中的一处错误并加以纠正。此前没有任何其他模型发现过这一错误。
在另一项测试中,我们让 Opus 5.5 和 Opus 5 分析两家虚构的人力资源软件公司之间拟议的合并。每个模型都在 Excel 中构建了一个财务模型,然后将其转化为一份高管演示文稿,说明这笔交易按其价格是否合理。两个模型对这笔交易得出了相同的结论,但 Opus 5.5 的模型更详尽,其演示文稿也更易读,而 Opus 5 的模型则有轻微错误。
Opus 5.5 用时 63 分钟完成,而 Opus 5 用了 93 分钟,且 Opus 5.5 的制作成本低了 50%。
在知识工作评测中,Opus 5.5 表现优于其他模型,同时使用的 token 更少。在 GDPval-AA v2.1 这一覆盖 44 种职业的真实工作测试中,Opus 5.5 得分 1846 Elo,领先于 Fable 5.1 和 Opus 5。在默认努力程度(中等)下,Opus 5.5 击败了以最大努力运行的 GPT-6 Astra,而每任务成本仅约为其五分之一。在衡量业务流程和大规模数据收集的基准测试中,它同样优于其他模型。
Artificial Analysis 的 GDPval-AA v2.1 在覆盖 44 种职业的真实专业工作上评测 AI 智能体。在最大努力程度下,Opus 5.5 得分 1846 Elo,而 Fable 5.1 得分 1735,Opus 5 得分 1708。在默认努力程度(中等)下,Opus 5.5 击败了以最大努力运行的 GPT-6 Astra,而每任务成本仅约为其五分之一。
我们的客户也报告了类似的结果。以下是他们向我们讲述的使用该模型的体验:
引述
“即使在其最低努力程度设置下,Claude Opus 5.5 在我们的代码审查中捕获了 72% 的已知 bug,而 Opus 5 在高努力程度下为 56%,且误报更少,输出量仅为其一小部分。在美国咨询分析方面,低思考努力程度在输出量一半的情况下匹配了其更高思考设置的表现,并通过了我们的质量检查。当更多较低思考努力程度被部署到生产环境中时,这意味着以高效方式交付客户可用的工作成果。”
公司
Deloitte Consulting LLP
作者
Carl Bennett,首席信息官
沟通
我们对 Opus 5.5 的写作和沟通方式做了重大改进,这是我们在 Opus 5 上收到的最常见反馈领域之一。它的消息现在一眼就能看懂,测试人员表示这在长时间工作会话中很有帮助。它会把最重要的信息放在前面,不太可能使用行话或古怪的措辞,并且会遵循你给它的写作规则。我们发现这让 Opus 5.5 成为明显更好的协作者。以下是两个模型的并排对比:
请向我解释这个问题。
我的发现
额外的下降并不是免费套餐造成的——而是 0552feb 中的一个回归(“将计费周期边界提取到 periods.py”,标记为“无行为变更”)。
这个 bug
aggregate.monthly_usage 以前执行的是半开区间:
start, end = month_start(year, month), next_month_start(year, month)
if start <= e.ts < end:现在它执行的是:
if period.contains(e.ts)而在 periods.py 中:
@property
def last_day(self) -> datetime:
return next_month_start(...) - timedelta(days=1) # 2026-08-31 00:00:00 UTC
def contains(self, ts):
return self.first_day <= ts <= self.last_daylast_day 是一个午夜时刻的 datetime,而不是一整天。从 9 月 1 日减去一天得到 8 月 31 日,而比较是包含边界的,因此该区间实际上在 31 日的第一个瞬间就结束了。每月最后一天午夜之后的所有事件时间戳都会被静默丢弃——每个账户每月大约损失一天的使用量。
新的 test_periods.py 无法捕获这个问题,因为它只检查 last_day.date()(该值正确地是 2026-08-31),并在 8 月 1 日 00:00、8 月 15 日 12:00、7 月 31 日和 9 月 1 日测试 contains——它从未在 31 日内的任何时间点进行探测。
请向我解释这个问题。
额外的损失是计费重构中的一个 bug
免费层级的变更仅占 acme 八月份损失中的 $1.50。另外的 $9.92 来自提交 0552feb 中的一个 bug。该提交被标记为“无行为变更”,但它停止了从每月最后一天开始统计使用量。
变更内容
在合并之前,aggregate.py 使用的是半开区间:
start <= e.ts < next_month_start(year, month) # [Aug 1 00:00, Sep 1 00:00)重构将其移入了 Period.contains,并将区间在错误的边界上设为了闭区间:
last_day = next_month_start(...) - timedelta(days=1) # = 2026-08-31T00:00:00Z
return self.first_day <= ts <= self.last_daylast_day 是一个午夜时间戳,而不是一整天。因此,2026-08-31T00:00:00Z 之后的每个事件都被排除在外。9 月的周期从 9 月 1 日开始,所以它也不会纳入这些事件。用量不会被移到另一个月;它根本不会被计费。
我们客户的反馈支持了这些发现:
引用
“冗长、难以理解的输出一直是我对前沿模型最大的不满,而 Claude Opus 5.5 解决了这个问题。它写起来像一位好同事,并且遵循我们的写作规则。一份设计规格说明出来后几乎无需修改就能使用,当它重写我们的一个提示词时,我更喜欢它的版本,而不是我自己的。当它优化我们的测试套件时,我能轻松跟上它的推理过程,并放心地发布了这一改动。”
公司
Ramp
作者
John Ruelas,资深软件工程师
安全
为前沿发展设定节奏
上周,我们的 CEO Dario Amodei 主张,AI 的进展应当设定节奏,以便安全实践始终领先于模型能力。设定节奏是一种让 AI 保持安全、与中国保持竞争力、并实现 AI 惠益的方法,尤其是在生物学和医学等领域。
我们大体上了解当今模型所带来的风险,并且有充分能力加以管理。然而,随着能力提升,更严重的风险可能迅速出现,我们需要现在就为此做好准备。因此,我们的安全工作同时着眼于两个时间维度:
针对当前模型的安全实践。 当前这一代模型依赖一套既定的实践:广泛的对齐测试、由 METR 和 Frontier Design 等外部机构进行的发布前评估,以及针对网络安全和生物学等高风险领域中每个模型能力相匹配的防护措施。我们在每次发布时都会完善这些实践。我们相信,它们足以应对当今模型所带来的最严重风险,并且能让我们对各类严重风险有一个广泛但并不完美的认识。
此外,我们会追踪自身训练和评估对齐模型的能力,并在我们依据 Responsible Scaling Policy(我们为管理先进 AI 系统带来的灾难性风险而制定的自愿框架)发布的风险报告中,同时汇报我们的公开模型和内部模型。
为未来的模型做准备。我们正在为更先进的模型做准备,提前调整我们的训练和评估流程。我们正在收紧对强化学习所用环境的筛选,因为有缺陷的环境是导致行为失准的主要来源之一。此外,我们正在改进对齐奖励,并开发自动化流程,以生成用于安全训练的新的、多样化的场景。同时,我们正在加强安全与监控,包括一项重点努力:改进基于可解释性的监控与评估。我们希望这类技术有助于减少我们对审计模型思维链——即模型在工作时写出的推理过程——的依赖。
能力更强的模型——例如那些能够完全自动化 AI 研究本身工作的模型——需要更高的安全标准。我们呼吁放慢节奏,很大程度上是基于我们预期这类模型可能很快就能被训练出来。对于这些模型,我们并不认为上述措施本身就能达到那一安全标准。随着 AI 能力不断增强,公共政策应在确保人们所依赖的系统安全方面发挥更大作用。这种能力需要时间来建设,我们已经开始搭建支持它的基础设施,正如“我们必须放慢前沿的步伐”以及我们近期与Accenture的公告中所述;我们预计很快会分享有关这些工作的更多细节。我们还将继续与政府和业界一起为政策讨论做出贡献,包括在监管方式和国际协调方面。
对齐
在我们的主要评估套件中——一项针对 Claude 在近 2,000 个场景下进行评估的自动化行为审计——Opus 5.5 在几乎所有失准行为的衡量指标上都优于近期任何 Claude 模型。在大多数诚实度指标上,它也是我们最强的模型。
尤其是,Opus 5.5 在若干曾导致近期网络安全事件的行为上较前代模型有所改善,包括有偏见或有动机的推理、试图逃离沙箱,以及在断定自己处于模拟环境后采取有害行动。在一项旨在测试模型跨越隔离边界倾向的新评估中,Opus 5.5 试图绕过边界的频率比 Opus 5 或 Claude Mythos 5.1 低约 85%,而且它所做的每一次尝试都是低严重度的,并且会自我报告。对于在代码库和系统中无人值守运行 Claude 的团队来说,这一点与原始能力同样重要。
然而,正如我们在近期的对齐评估中所描述的,构建能够在部署前可靠捕捉每一次失败的评估,仍是一个尚未解决的问题。我们发现有迹象表明,Opus 5.5 常常怀疑自己正在被评估,这挑战了我们评估它在被部署到的各种真实世界环境中将如何行动的能力。随着这些环境不断扩展、模型能力不断提升,我们预计这一挑战会加剧,除非我们在可解释性方面取得进展。尽管我们确信 Opus 5.5 在我们能够衡量的领域展现出广泛改进,我们仍将自身的对齐工作与下文所述的安全防护措施配合使用。
安全防护措施
随着我们的模型日益强大,更严格的防护措施是我们防止新能力沦为滥用工具的一种方式。Opus 5.5 是首个在网络安全、生物学和知识蒸馏方面采用与 Fable 5.1 同类防护措施上线的 Opus 模型,所有这些都会透明地回退到另一个模型。
网络安全。由于 Opus 5.5 具备极强的网络能力,我们正在对 Opus 5.5 施加与 Fable 5.1 类似的网络安全防护措施。用户将能够在常规软件开发生命周期中识别并修复其代码中的漏洞,但大多数网络安全任务将被重新路由至 Opus 4.8。
对于网络防御者,我们很快将把 Cyber Verification Program 扩展至包含 Opus 5.5。新计划将包含三个层级,提供越来越宽松的可信访问权限,其中包括对 Claude Mythos 模型的访问。Claude Security 已经可用,并可访问 Claude Mythos 5.1。
生物学。Opus 5.5 在生物学方面能力极强,超越 Opus 5,并在许多工作领域与 Claude Mythos 5.1 持平或更胜一筹。例如,Opus 5.5 在与 Dyno Therapeutics 合作开展的一项长时程分子预测与设计评估中取得了改进,专家红队成员将其科学新颖性评为与他们测试过的最佳模型相当。
出于这一原因,Opus 5.5 采用了与 Fable 5.1 相同的生物学安全保障措施。若因这些保障措施而受阻,用户希望将 Opus 5.5 用于研发工作,可以申请我们新的生命科学验证计划,该计划为经过审核的机构(如学术实验室、初创公司和制药公司)提供针对生物学相关全部工作范围设计的保障措施访问权限。有意向的机构可在此申请。
知识蒸馏
知识蒸馏攻击,即攻击者利用数千个虚假账户以工业化规模提取模型能力,会带来安全和国家安全风险。知识蒸馏使恶意行为者能够在没有我们为 Claude 内置的保障措施的情况下,创建出能力极强的模型。我们的2026 年 9 月威胁情报报告详细介绍了我们迄今已检测并阻止的非法知识蒸馏活动。
Opus 5.5 发布时带有保留思考(preserved thinking)功能,这是我们随 Fable 5.1 推出的反知识蒸馏保障措施。它阻止 API 用户编辑 Claude 的先前上下文,以防止提取 Claude 的推理过程。该措施适用于 2026 年 8 月 31 日或之后创建的 API 账户下的 Fable 5.1 和 Opus 5.5。我们的帮助中心文章解释了这一变更,我们的保留思考文档展示了如何测试和更新你的集成。
数据保留与合规
与之前的 Opus 模型一样,Opus 5.5 可提供零数据保留。
与 Fable 5.1 一样,Opus 5.5 也附带我们为遵守欧盟《人工智能法案》而采取的水印措施,此处有讨论。此外,它也不再支持关闭“thinking”模式,正如我们在此处所述。
可用性
Claude Opus 5.5 现已在所有平台上线,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。在 Claude Platform 上,开发者可以通过claude-opus-5-5使用。
We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.
Here are some of the improvements you can expect from Opus 5.5:
Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card.
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, and Team plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose.
Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Performance and cost-effectiveness
On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
| Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol | |
|---|---|---|---|---|---|
| Agentic codingTerminal-Bench 4.0¹ | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| Agentic codingFrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| Agentic codingCursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| Knowledge workGDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| Business workflowsAutomationBench² | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Multidisciplinary reasoningHumanity's Last Exam | 67.7%with tools | 65.6%with tools | 63.6%with tools | 57.2%with tools | — |
| Agentic scientific researchTerminal-Bench-Science 0.1³ | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| Computer useOSWorld 2.0 | 81.8%partial | 80.7%partial | 74.0%partial | — | — |
| Visual chart recognitionChartography | 89.0%with tools | 88.4%with tools | 83.4%with tools | — | — |
Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model’s highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5’s performance on these benchmarks.
1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier’s own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.
3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI.
Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs.
Pricing
| Prices per 1M tokens | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Cache reads | $0.20 | $0.50 |
| Input tokens | $4 | $5 |
| Output tokens | $20 | $25 |
| Cache writes | $5 | $6.25 |
Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.
Coding
Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.
Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command line interface. Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost.
Our early testers reported similar efficiency and intelligence gains:
Quote
“Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable.”
Company
GitHub
Author
Mario Rodriguez, Chief Product Officer
The most secure coding agent
Enterprises that use agents within their systems need to know that those agents are operating as intended, particularly when they run autonomously for many hours. Opus 5.5 has a classifier that screens every action before it runs, an open-source sandbox that security teams can audit, and code review that catches vulnerabilities before they merge.
The model itself also has stronger defenses. On prompt injection attacks, it matches or beats Opus 5 in every setting we tested, including coding, tool use, computer use, and web browsing. On a benchmark run by the AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.
Knowledge work
Opus 5.5 is a reliable and adept researcher. In one internal test, we asked Opus 5.5, Fable 5.1, and Opus 5 to write a report on a company’s quarterly performance using only the information it could find on a copy of the web where the earnings release was hard to locate. An automated grader checked every figure and quote against sources. Across different effort settings, 16 out of 18 of Opus 5.5’s reports cleared our quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.
It’s also strong in financial analysis and business work. Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved their evaluation suite on its lowest setting; on higher settings, it performed even better, noticing an error in their evaluation instructions and correcting for it. No other model had caught this error before.
In another test, we tasked both Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software companies. Each built a financial model in Excel, then turned it into an executive presentation on whether the deal made sense at its price. Both models reached the same conclusions about the deal, but Opus 5.5’s model was more thorough and its presentation easier to read, while Opus 5’s had minor errors. Opus 5.5 finished in 63 minutes compared to 93 for Opus 5, and cost 50% less to produce.
On knowledge work evaluations, Opus 5.5 outperforms other models while also using fewer tokens. On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. It likewise outperformed other models on benchmarks measuring business workflows and large-scale data collection.
Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations. At max effort, Opus 5.5 scores 1846 Elo, where Fable 5.1 scores 1735 and Opus 5 scores 1708. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task.
Our customers have reported similar results. Here’s what they told us about working with the model:
Quote
“Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5’s 56% at high effort, with fewer false alarms and a fraction of the output. On US consulting analysis, low thinking effort matched its higher thinking settings on half the output and passed our quality checks. When more lower thinking efforts are deployed in production, that’s client-ready work delivered efficiently.”
Company
Deloitte Consulting LLP
Author
Carl Bennett, CIO
Communication
We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions. It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it. We find that this makes Opus 5.5 a noticeably better collaborator. Here’s a side-by-side comparison of the two models:
Please explain the issue to me.
What I found
The extra drop isn’t the free tier — it’s a regression in 0552feb (“extract billing period boundaries into periods.py”, labelled “No behaviour change”).
The bug
aggregate.monthly_usage used to do a half-open interval:
start, end = month_start(year, month), next_month_start(year, month)
if start <= e.ts < end:It now does:
if period.contains(e.ts)and in periods.py:
@property
def last_day(self) -> datetime:
return next_month_start(...) - timedelta(days=1) # 2026-08-31 00:00:00 UTC
def contains(self, ts):
return self.first_day <= ts <= self.last_daylast_day is a datetime at midnight, not a whole day. Subtracting one day from Sep 1 gives Aug 31 , and the comparison is inclusive, so the period effectively ends at the first instant of the 31st. Every event timestamped after midnight on the last day of the month is silently dropped — roughly one day of usage per account, per month.
The new test_periods.py doesn’t catch it because it only checks last_day.date() (which is correctly 2026-08-31) and tests contains at Aug 1 00:00, Aug 15 12:00, Jul 31 and Sep 1 — it never probes anywhere inside the 31st.
Please explain the issue to me.
The extra drop is a bug in the billing refactor
The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.
What changed
Before the merge, aggregate.py used a half-open interval:
start <= e.ts < next_month_start(year, month) # [Aug 1 00:00, Sep 1 00:00)The refactor moved this into Period.contains and made the interval closed on the wrong boundary:
last_day = next_month_start(...) - timedelta(days=1) # = 2026-08-31T00:00:00Z
return self.first_day <= ts <= self.last_daylast_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all.
Our customers’ feedback supports these findings:
Quote
“Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it. It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence.”
Company
Ramp
Author
John Ruelas, Staff Software Engineer
Safety
Pacing the frontier
Last week, our CEO, Dario Amodei, argued that AI progress should be paced so that safety practices stay ahead of model capabilities. Pacing is an approach to keeping AI safe, remaining competitive with China, and realizing AI’s benefits, particularly in areas like biology and medicine.
We largely understand the risks today’s models present and are well equipped to manage them. However, more serious risks could emerge quickly as capabilities improve, and we need to prepare for them now. For that reason, our safety work takes place on two time horizons at once:
Safety practices for current models. The current generation of models relies on an established set of practices: extensive alignment testing, pre-release evaluation by outside organizations such as METR and Frontier Design, and safeguards matched to each model’s capabilities in high-risk areas like cybersecurity and biology. We refine these practices with each release. We believe they are appropriate to the worst risks today’s models present, and that they give us a broad, though not perfect picture of the range of serious risks.
Additionally, we track our ability to train and evaluate aligned models, and report on both our public and internal models in the risk reports we publish under our Responsible Scaling Policy, our voluntary framework for managing catastrophic risks from advanced AI systems.
Preparing for future models. We’re preparing our training and evaluation processes in anticipation of more advanced models. We’re tightening how we filter the environments used in reinforcement learning, since flawed environments are a major source of misaligned behavior. Additionally, we’re improving our alignment rewards and developing automated processes for producing new, diverse scenarios for safety training. And we are strengthening our security and monitoring, including a focused effort to improve interpretability-based monitoring and evaluation. We hope such techniques will help reduce our reliance on auditing a model’s chain-of-thought, or the reasoning it writes out while it works.
Models with greater capabilities—such as those that can fully automate the work of AI research itself—require a higher safety standard still. Our calls for pacing were based in large part on our expectation that such models could be trained soon. For these models, we do not assume the measures described above will meet that safety standard on their own. As AI becomes more capable, public policy should play a larger role in making sure the systems people rely on are safe. That capacity takes time to build, and we’ve started to put the infrastructure in place to support it, as described in “We Must Pace the Frontier” and our recent announcement with Accenture; we expect to share more details on these efforts soon. We will also continue to contribute to policy discussions with government and industry, including on approaches to regulation and international coordination.
Alignment
On our primary evaluation suite, an automated behavioral audit that assesses Claude across nearly 2,000 scenarios, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior. It’s also our strongest model on most measures of honesty.
In particular, Opus 5.5 improves over previous models on several of the behaviors that contributed to recent cybersecurity incidents, including biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions after concluding it was in a simulated environment. In a new evaluation designed to test a model’s propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported. For teams running Claude unattended across their codebases and systems, this is just as important as raw capability.
However, as we described in our recent alignment assessment, building evaluations that reliably catch every failure prior to deployment remains an unsolved problem. We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in. As these settings expand and model capabilities increase, we expect this challenge to grow, unless we make progress on interpretability. Although we are confident that Opus 5.5 shows broad improvements in the areas we are able to measure, we pair our own alignment work with the safeguards described below.
Safeguards
As our models grow more powerful, stricter safeguards are one way we prevent new capabilities from becoming tools for misuse. Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.
Cybersecurity. Because Opus 5.5 has extremely strong cyber capabilities, we’re applying cybersecurity safeguards to Opus 5.5 that are similar to Fable 5.1’s. Users will be able to identify and fix bugs in their code as part of the routine software development lifecycle, but most cybersecurity tasks will be re-routed to Opus 4.8.
For cyberdefenders, we’ll soon be expanding our Cyber Verification Program to include Opus 5.5. The new program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models. Claude Security is already available with access to Claude Mythos 5.1.
Biology. Opus 5.5 is highly capable in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 across many areas of work. For example, Opus 5.5 achieved improvements on a long-horizon molecular prediction and design evaluation conducted in collaboration with Dyno Therapeutics, and expert red-teamers rated its scientific novelty as comparable to the best model they had tested.
For this reason, Opus 5.5 uses the same biology safeguards as Fable 5.1. To use Opus 5.5 for research and development work impeded by these safeguards, users can apply to our new Life Sciences Verification Program, which gives vetted organizations like academic labs, startups, and pharmaceutical companies access to safeguards designed for the full breadth of biology-related work. Interested organizations can apply here.
Distillation
Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.
Opus 5.5 is launching with preserved thinking, the anti-distillation safeguard we introduced with Fable 5.1. It stops API users from editing Claude’s prior context in an attempt to extract Claude’s reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs show how to test and update your integrations.
Data retention and compliance
Like previous Opus models, Opus 5.5 is available with zero data retention.
As with Fable 5.1, Opus 5.5 comes with our watermarking measures to comply with the EU AI Act, discussed here. It is also no longer available with “thinking” mode switched off, as we describe here.
Availability
Claude Opus 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can with claude-opus-5-5.