跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前精选AI 评分83
AI 导读

Anthropic 发布 Claude Sonnet 5.5,Terminal-Bench 4.0 得分 70.6%,远高于 Sonnet 5 的 10.3%,价格维持 $2/$10 每百万输入/输出 token。

推荐理由

原文对比了 Sonnet 5.5 的跑分和每任务成本,指出低成本档位也能超过上代最高分,阅读时可关注定价经济性这条主线。

正文 · 原文

Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices.

Overall, 30% cost reduction per-task due to faster speeds and fewer tool calls.

Keeps Sonnet 5's $2/$10 per million input/output tokens, half Opus 5.5's rates.

Its savings instead come from doing less work per job, since Anthropic says fewer tokens and tool calls cut the total cost of a task by up to 30%.

Output also arrives more than 30% faster than Sonnet 5's, making Sonnet 5.5 Anthropic's quickest Sonnet yet.

Users can dial an effort setting, trading longer reasoning and more self-checking for a higher cost per task.

The economics might count for more than the leaderboard. Sonnet 5.5 operating at Low or Medium effort is able to exceed Sonnet 5's top score at about one-tenth the cost per task. On FrontierCode, it says Sonnet 5.5 at High effort scores roughly 10 points above Sonnet 5 at the same setting while costing approximately one-fifteenth as much per task.

Anthropic characterizes Sonnet 5.5 as ideal for comparatively well-defined routine work, such as software debugging, coding, document production, building presentations and spreadsheets, and designing or refining interfaces.

引用Claude@claudeai
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
在 X 查看被引用的帖子

来源:Rohan Paul · x.com