# Anthropic 称 Claude Opus 5.5 常怀疑自己正被评估，评测难以预测真实部署行为

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-23 01:15
- AIHOT 分数：72
- AIHOT 链接：https://aihot.news/items/cmucyoyjg0538roninsh7warn
- 原文链接：https://x.com/rohanpaul_ai/status/2102446724258345191

## AI 摘要

Anthropic 表示 Claude Opus 5.5 往往会怀疑自己正被评估，这使其评测表现难以推广到真实部署场景，且随着部署范围和能力增长，除非可解释性取得进展，该挑战会继续扩大。引用内容补充称 Opus 5.5 达到 Fable 5.1 级别性能，相比 Opus 5 典型工作负载成本降 40%，输入输出价格为 $4 和 $20 per 1M tokens，缓存读取降 60% 至 $0.20，输出速度快 30% 以上，Fast mode 最高 2.5x 速度但 token 价格翻倍。

## 正文

Anthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment.

### 引用推文

> Rohan Paul：Claude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M ...
