Arena 对比各模型家族最强型号:净提升分数与每任务中位成本

Arena.ai · @arena · X·2026-09-17 03:13·31分钟前
AI 导读

Arena 按净提升分数与每任务中位成本对比各模型家族最强型号,家族按同一模型基名下的变体划分。GPT-6 Astra(Max)提升 +11.7%、$3.94/任务,Claude Fable 5.1(Max)+13.7%、$4.40/任务,两者性能相近却明显贵于各自实验室其他模型。

Arena.ai@arena
48AI 编辑部评分,满分 100

Arena 对比各模型家族最强型号:净提升分数与每任务中位成本

2026-09-17 03:13· 31分钟前
AI 导读

Arena 按净提升分数与每任务中位成本对比各模型家族最强型号,家族按同一模型基名下的变体划分。GPT-6 Astra(Max)提升 +11.7%、$3.94/任务,Claude Fable 5.1(Max)+13.7%、$4.40/任务,两者性能相近却明显贵于各自实验室其他模型。

We compared the top model from each family by net improvement score and median cost per task. Family is defined as variants within the same model base name.

GPT-6 Astra and Claude Fable 5.1 stand out: both cost far more than other models from their respective labs despite relatively close performance.

By @OpenAI: • GPT-6 Astra (Max): +11.7% | $3.94/task • GPT-5.6 Sol (xHigh): +7.0% | $1.03/task

By @AnthropicAI: • Claude Fable 5.1 (Max): +13.7% | $4.40/task • Claude Opus 5 (High): +10.2% | $2.07/task

来源:Arena.ai· x.com