🚨 AI News | TestingCatalog · @testingcatalog · X·2026-09-23 00:39·1小时前
AI 导读

Anthropic 发布 Claude Opus 5.5,称其为 Claude 5.5 系列首款模型,大多数任务表现达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。官方基准显示其在 Terminal-Bench 4.0 智能体编码得分 66.4%,OSWorld 2.0 计算机使用得分 81.8%,测试采用自适应思考最高推理力度并启用生产环境安全防护。

🚨 AI News | TestingCatalog@testingcatalog
73AI 编辑部评分,满分 100
2026-09-23 00:39· 1小时前
AI 导读

Anthropic 发布 Claude Opus 5.5,称其为 Claude 5.5 系列首款模型,大多数任务表现达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。官方基准显示其在 Terminal-Bench 4.0 智能体编码得分 66.4%,OSWorld 2.0 计算机使用得分 81.8%,测试采用自适应思考最高推理力度并启用生产环境安全防护。

BREAKING 🔥: Claude Opus 5.5 has been released and it is a new SOTA model in agentic coding tasks and computer use!

Exponential slowdown 👀

ClaudeIntroducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to ru...