Anthropic 正在推出 Claude Opus 5.5,这是全新模型家族中的首个模型。该公司表示,它能提供 Claude Fable 5.1 级别的性能,同时成本显著更低,运行速度也比其前代更快。
据 Anthropic 称,Opus 5.5 在"大多数任务上"与 Claude Fable 5.1 相当,而运行成本比 Opus 5 低约 40%。Claude Sonnet 5.5 和 Haiku 5.5 预计将在未来几周内推出,在性能、效率和安全方面具有类似的提升。Anthropic 的基准测试显示,新的 Opus 模型在大多数任务上领先于 Fable 5.1 和 OpenAI 价格高得多的 GPT-6 Astra。
| 基准测试 / 能力 | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| 智能体编程 Terminal-Bench 4.0[1] | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| 智能体编程 FrontierCode v1.1(主) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| 智能体编程 CursorBench 4.0 | 57.8% | 51.8% | 46.6% | N/A | 41.7% |
| 知识工作 GDPval-AA v2.1 | 1,846 | 1,735 | 1,708 | 1,542 | 1,588 |
| 业务流程 AutomationBench[1] | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| 多学科推理 Humanity's Last Exam | 67.7% (使用工具) | 65.6% (使用工具) | 63.6% (使用工具) | 57.2% (使用工具) | 不适用 |
| 智能体科学研究 Terminal-Bench-Science 0.1[1] | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| 计算机使用 OSWorld 2.0 | 81.8% (部分) | 80.7% (部分) | 74.0% (部分) | 不适用 | 不适用 |
| 可视化图表识别 Chartography | 89.0% (使用工具) | 88.4% (使用工具) | 83.4% (使用工具) | 不适用 | 不适用 |
Anthropic 表示,新系列主要回应了客户在成本、效率和沟通质量方面的反馈,尤其是在金融服务、法律和软件开发领域。
更低的价格和 token 用量降低了运营成本
Anthropic 将 Opus 5.5 定价为每百万输入 token 4 美元、每百万输出 token 20 美元,低于 Opus 5 的 5 美元和 25 美元。这意味着 token 价格下降了 20%。该公司还将缓存读取成本降低了 60%。
Anthropic 表示,总运营成本(同时计入 token 价格和 token 用量)应比 Opus 5 低约 40%。该模型使用的 token 更少,输出生成速度提升超过 30%。
| 每 1M token 价格 | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| 缓存读取 | $0.20 | $0.50 |
| 输入 token | $4 | $5 |
| 输出 token | $20 | $25 |
| 缓存写入 | $5 | $6.25 |
订阅用户的五小时使用限额将提高 20%。Anthropic 表示,得益于该模型更低的成本,这些限额总体上还能再延长 25%。用户还可以保留一次限额重置,留到最需要的时候使用。
Anthropic 正借助编程基准测试来论证其在价格和性能上的优势。在 FrontierCode 上,该公司称 Opus 5.5 击败了 OpenAI 的 GPT-6 Astra,而每任务成本仅约为后者的 20%。在 Terminal-Bench 4.0 上,它声称以 40% 的成本实现了与 Astra 相同的性能。在 CursorBench 上,它表示 Opus 5.5 以三分之一的成本领先 GPT-5.6 Sol 11 分。
此次降价是对来自 OpenAI、尤其是中国 AI 模型压力的回应,这些模型性能较低,但成本仅为前者的一小部分。
Anthropic 承诺减少“Claude 腔”
Opus 5.5 还应当比早先的模型表达得更自然。Anthropic 表示,它会把最重要的信息放在最前面,少用行话,并且更严格地遵循写作指令。早期测试者称其文字更清晰、更易懂,Anthropic 表示这使它成为长时间工作会话中更好的伙伴。当前的 Claude 模型因其公式化、晦涩难懂的文风,有时被称为“Claude 腔”,而饱受批评。

Opus 5.5 也是首个在网络安全、生物学和前沿 LLM 开发方面拥有与 Fable 5.1 同等防护措施的 Opus 模型。Anthropic 表示,当这些防护措施触发时,请求会被透明地路由到另一个模型。
用户仍然可以在自己的代码中发现并修复 bug,但大多数网络安全任务将交由较旧的 Opus 4.8 处理。被分类器标记为涉及生物学或前沿 LLM 开发的请求将交由 Opus 5 处理。
经过验证的组织可以通过生命科学验证计划申请将该模型用于生物学研究。Anthropic 计划在未来几周内将其现有的网络验证计划扩展到 Opus 5.5。
Claude Opus 5.5 现已在所有平台上线,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。使用 Claude Platform 的开发者可以通过模型 ID claude-opus-5-5 访问它。
Anthropic 呼吁在模型能力日益增强之际提高安全标准
Anthropic 计划更严格地审查强化学习环境。该公司表示,有缺陷的训练环境是模型行为失准的主要来源。它还在致力于更好的对齐奖励、自动创建安全训练场景的方法,以及更强的安全和监控措施。
AI 实验室正在争论应以多快的速度发布能力更强的模型。OpenAI 最近呼吁为能够自我改进的 AI 系统制定国际标准,而多位研究人员已公开敦促各实验室在对齐方法跟上之前放缓发布。Anthropic 采取了类似立场,呼吁对可能完全自动化 AI 研究的模型制定更高的安全标准。该公司表示,公共政策应在制定该标准方面发挥更大作用。外部机构Frontier Design和METR在 Opus 5.5 发布前对其进行了测试。
新限制针对知识蒸馏,并支持欧盟 AI 法案合规
Anthropic 还针对其所谓的蒸馏攻击推出了应对措施。该公司表示,攻击者利用数千个虚假账号以工业化规模提取模型能力,并在没有其安全保障的情况下构建出能力极强的模型。Anthropic 援引了一份 2026 年 9 月的威胁报告,其中记录了其迄今已检测并阻止的非法蒸馏活动。
Opus 5.5 发布时搭载了“Preserved Thinking”,这是一项最早随 Fable 5.1 推出的反蒸馏措施。它阻止 API 用户编辑 Claude 此前的上下文以提取其推理过程。该措施适用于 2026 年 8 月 31 日或之后创建的 API 账号下的 Fable 5.1 和 Opus 5.5。
Opus 5.5 还包含水印措施,以符合欧盟《人工智能法案》。该模型不再能在禁用“Thinking”模式的情况下运行,并且支持零数据保留。
Anthropic is launching Claude Opus 5.5, the first model in a new family. The company says it delivers Claude Fable 5.1-level performance while costing significantly less and running faster than its predecessor.
According to Anthropic, Opus 5.5 matches Claude Fable 5.1 "on most tasks" while costing about 40 percent less to run than Opus 5. Claude Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks, with similar gains in performance, efficiency, and safety. Anthropic's benchmarks show the new Opus model ahead of both Fable 5.1 and OpenAI's much more expensive GPT-6 Astra on most tasks.
| Benchmark / capability | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Agentic coding Terminal-Bench 4.0[1] | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| Agentic coding FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| Agentic coding CursorBench 4.0 | 57.8% | 51.8% | 46.6% | N/A | 41.7% |
| Knowledge work GDPval-AA v2.1 | 1,846 | 1,735 | 1,708 | 1,542 | 1,588 |
| Business workflows AutomationBench[1] | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Multidisciplinary reasoning Humanity's Last Exam | 67.7% (with tools) | 65.6% (with tools) | 63.6% (with tools) | 57.2% (with tools) | N/A |
| Agentic scientific research Terminal-Bench-Science 0.1[1] | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| Computer use OSWorld 2.0 | 81.8% (partial) | 80.7% (partial) | 74.0% (partial) | N/A | N/A |
| Visual chart recognition Chartography | 89.0% (with tools) | 88.4% (with tools) | 83.4% (with tools) | N/A | N/A |
Anthropic says the new series primarily addresses customer feedback on cost, efficiency, and communication quality, particularly in financial services, law, and software development.
Lower prices and token usage cut operating costs
Anthropic prices Opus 5.5 at $4 per million input tokens and $20 per million output tokens, down from $5 and $25, respectively, for Opus 5. That's a 20 percent cut in token prices. The company also reduced cache read costs by 60 percent.
Anthropic says total operating costs, which account for both token prices and token usage, should be about 40 percent lower than Opus 5's. The model uses fewer tokens and generates output more than 30 percent faster.
| Prices per 1M tokens | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Cache reads | $0.20 | $0.50 |
| Input tokens | $4 | $5 |
| Output tokens | $20 | $25 |
| Cache writes | $5 | $6.25 |
Five-hour usage limits for subscribers will increase by 20 percent. With the model's lower costs, Anthropic says those limits stretch 25 percent further overall. Users can also save a limit reset for when they need it most.
Anthropic is using coding benchmarks to make its case on price and performance. On FrontierCode, the company says Opus 5.5 beats OpenAI's GPT-6 Astra at about 20 percent of the cost per task. On Terminal-Bench 4.0, it claims the same performance as Astra at 40 percent of the cost. On CursorBench, it says Opus 5.5 beats GPT-5.6 Sol by 11 points at one-third of the cost.
The price cut is a response to pressure from OpenAI and especially Chinese AI models, which offer lower performance but cost a fraction as much.
Anthropic promises less "Claudish"
Opus 5.5 is also supposed to communicate more naturally than earlier models. Anthropic says it puts the most important information first, uses less jargon, and follows writing instructions more closely. Early testers described its writing as clearer and easier to understand, which Anthropic says makes it a better partner for long work sessions. Current Claude models have drawn plenty of criticism for their formulaic, convoluted writing, sometimes called "Claudish".

Opus 5.5 is also the first Opus model with safeguards for cybersecurity, biology, and frontier LLM development that match those of Fable 5.1. When those safeguards kick in, Anthropic says requests are transparently routed to another model.
Users can still find and fix bugs in their code, but most cybersecurity tasks will go to the older Opus 4.8. Requests flagged by classifiers for biology or frontier LLM development will go to Opus 5.
Verified organizations can apply to use the model for biological research through the Life Sciences Verification Program. Anthropic plans to extend its existing Cyber Verification Program to Opus 5.5 in the coming weeks.
Claude Opus 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers using the Claude Platform can access it with the model ID claude-opus-5-5.
Anthropic calls for higher safety standards as models become more capable
Anthropic plans to screen reinforcement learning environments more strictly. The company says flawed training environments are a major source of misaligned model behavior. It's also working on better alignment rewards, automated ways to create safety training scenarios, and stronger safety and monitoring measures.
AI labs are debating how quickly to release more capable models. OpenAI recently called for international standards for AI systems that could improve themselves, while several researchers have publicly urged labs to slow releases until alignment methods catch up. Anthropic is taking a similar position, calling for a higher safety standard for models that could automate AI research entirely. The company says public policy should have a greater role in setting that standard. External organizations Frontier Design and METR tested Opus 5.5 before its release.
New restrictions target distillation and support EU AI Act compliance
Anthropic is also introducing measures against what it calls distillation attacks. The company says attackers use thousands of fake accounts to extract a model's capabilities at an industrial scale and build highly capable models without its safeguards. Anthropic cites a September 2026 threat report documenting illegal distillation activity it has detected and stopped so far.
Opus 5.5 launches with "Preserved Thinking," an anti-distillation measure first introduced with Fable 5.1. It prevents API users from editing Claude's prior context to extract its reasoning. The measure applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026.
Opus 5.5 also includes watermarking measures to comply with the EU AI Act. The model can no longer run with "Thinking" mode disabled and it is available with Zero Data Retention.