Arena 公布 Claude Sonnet 5.5 (Max) 在 Agent Arena 首秀排名 #3,净提升 +12.5%,比 Claude Sonnet 5 (High) 高 8.1 个百分点,并在 Chat 类别以 +15.6% 排名 #1。
同一新闻,精选展示《Arena 评测:Claude Sonnet 5.5 登顶 Agent Arena 第 3 名但未入 Pareto 前沿》
Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement!
This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%).
This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58.
@AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team!
Exciting update: Claude Sonnet 5.5 with xHigh reasoning has landed in the Code Arena: WebDev. With 1786 pts, its ranked #3! At a blended $8/M tokens, Claude Sonnet 5.5 remains on the Pareto frontier with xHigh reasoning. This release is just 2 pts from GPT-6 Astra in the #2 spot with 1788 pts, for 80% of the price. By domain, Claude Sonnet 5.5 (xHigh) landed: - #2 in Gaming, Reference-Based Design, and Brand & Marketing - #3 in Simulations - #4 in Content Creation Tools and Consumer Product - #6 in Data & Analytics Congrats again to @AnthropicAI on this release!在 X 查看被引用的帖子
来源:Arena.ai · x.com