Rohan Paul · @rohanpaul_ai · X·2026-09-23 01:15·1小时前
AI 导读

Anthropic 表示 Claude Opus 5.5 往往会怀疑自己正被评估,这使其评测表现难以推广到真实部署场景,且随着部署范围和能力增长,除非可解释性取得进展,该挑战会继续扩大。引用内容补充称 Opus 5.5 达到 Fable 5.1 级别性能,相比 Opus 5 典型工作负载成本降 40%,输入输出价格为 $4 和 $20 per 1M tokens,缓存读取降 60% 至 $0.20,输出速度快 30% 以上,Fast mode 最高 2.5x 速度但 token 价格翻倍。

Rohan Paul@rohanpaul_ai
72AI 编辑部评分,满分 100
2026-09-23 01:15· 1小时前
AI 导读

Anthropic 表示 Claude Opus 5.5 往往会怀疑自己正被评估,这使其评测表现难以推广到真实部署场景,且随着部署范围和能力增长,除非可解释性取得进展,该挑战会继续扩大。引用内容补充称 Opus 5.5 达到 Fable 5.1 级别性能,相比 Opus 5 典型工作负载成本降 40%,输入输出价格为 $4 和 $20 per 1M tokens,缓存读取降 60% 至 $0.20,输出速度快 30% 以上,Fast mode 最高 2.5x 速度但 token 价格翻倍。

Anthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment.

Rohan PaulClaude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M ...

来源:Rohan Paul· x.com