Arena 宣布 OpenAI 的 GPT-6.1 Sol (Max) 在 Agent Arena 排名第 5(+11.23%),中位任务成本 $0.56,并重塑了 Pareto 前沿。
官方榜单数据给出了 GPT-6.1 Sol (Max) 的排名与成本对比,读者可以据此评估它在智能体任务上的性价比位置。
激动人心的消息:@OpenAi 的 GPT-6.1 Sol (Max) 刚刚登陆 Agent Arena,位列第 5(+11.23%),并重塑了帕累托前沿!
在每任务中位成本为 $0.56 的情况下,它的性能与 GPT-6 Sol 和 GPT-6 Astra 相差不到 2 个百分点,而成本却大幅降低:
- 成本比 GPT-6 Sol 低 39%,同时得分高出 +1.52 分
- 成本比 GPT-6 Astra 低 81%,同时差距在 1.04 分以内
与以下模型相比,GPT-6.1 Sol 也以大幅更低的成本实现了前五名的性能:
- 成本比 Claude Fable 5.1 (Max) 低 88%,同时差距在 3.08 分以内(排名第 1)
- 成本比 Claude Opus 5.5 (High) 低 65%,同时差距在 2.59 分以内(排名第 2)
- 成本比 Claude Sonnet 5.5 (Max) 低 80%,同时差距在 1.29 分以内(排名第 3)
祝贺团队 @OpenAI 发布这一版本!
激动人心的消息:@OpenAI 的 GPT-6.1 Sol (Max) 刚刚在 Code Arena: WebDev 上以 1759 分位列第 3,并且以每百万 Token 8 美元的混合价格重塑了帕累托前沿! GPT-6.1 Sol (Max) 在成本效率上实现了明显提升:在相同价格下,它比 GPT-6 Sol (Max) 提高了 70 分。 它与 GPT-6 Astra (Max) 的差距在 30 分以内,而混合 Token 成本低了 80%;与 Claude Opus 5.5 (Max) 相差 59 分,而价格只有其 50%。在下方的帖子中查看 Code Arena: WebDev 帕累托前沿上的位置。 总体而言,GPT-6.1 Sol 相比 GPT-6 Sol 提升了 4 个排名!它在每个类别中也有所提升: - 消费产品:第 5 → 第 1 - 模拟:第 6 → 第 3 - 数据与分析:第 4 → 第 3 - 内容创作工具:第 4 → 第 3 - 游戏:第 6 → 第 4 - 基于参考的设计:第 6 → 第 4 - 品牌与营销:第 10 → 第 6 祝贺 @OpenAI 团队发布!
原文
Exciting news: GPT-6.1 Sol (Max) by @OpenAI just landed the Code Arena: WebDev at #3 with 1759 pts, and at a blended $8/MToken it reshapes the Pareto frontier! GPT-6.1 Sol (Max) marks a clear improvement in cost efficiency: it gained 70 points over GPT-6 Sol (Max) for the same price. It landed within 30 points of GPT-6 Astra (Max) at 80% lower blended token cost, and 59 points from Claude Opus 5.5 (Max) at 50% of the price. See position on the Pareto frontier for the Code Arena: WebDev in the post below. Overall, GPT-6.1 Sol improved from GPT-6 Sol by 4 rankings! It also improved in every category: - Consumer Product: #5 → #1 - Simulations: #6 → #3 - Data & Analytics: #4 → #3 - Content Creation Tools: #4 → #3 - Gaming: #6 → #4 - Reference-Based Design: #6 → #4 - Brand & Marketing: #10 → #6 Congrats to the @OpenAI team on the release!
来源:Arena.ai · x.com