Artificial Analysis 能力指数 v1.1 更新发布

Artificial Analysis · @ArtificialAnlys · X·2026-09-15 00:07·6小时前
AI 导读

Artificial Analysis 发布 Capability Indices v1.1,将 O*NET 职业任务映射到基准并按能力出现频率加权,覆盖金融会计、战略运营、法律、医疗、工程和经济六大指数,基于 Intelligence Index v4.3 评测并结合专业评测。

Artificial Analysis@ArtificialAnlys
44AI 编辑部评分,满分 100

Artificial Analysis 能力指数 v1.1 更新发布

2026-09-15 00:07· 6小时前
AI 导读

Artificial Analysis 发布 Capability Indices v1.1,将 O*NET 职业任务映射到基准并按能力出现频率加权,覆盖金融会计、战略运营、法律、医疗、工程和经济六大指数,基于 Intelligence Index v4.3 评测并结合专业评测。

Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations

The Capability Indices map tasks from O*NET occupations to benchmarks that represent them, weighting each benchmark by how often its capability appears across the work. We cover six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics

Key changes: ➤ Finance & Accounting: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Finance), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work

➤ Strategy & Ops: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work

➤ Legal: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work

➤ Healthcare & Medical: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Long-Context Reasoning added as a new capability sourced from MLCR-AA, Agentic Customer Interaction removed from capabilities, and AA-Briefcase added to Agentic Knowledge Work

➤ Engineering: Terminal-Bench updated to v4.0 in Agentic Terminal Use, GPQA Diamond removed from Reasoning, and AA-Briefcase added to Agentic Knowledge Work

➤ Economics: AA-Briefcase added to Agentic Knowledge Work

Key results: ➤ Claude Fable 5.1 (max) leads in all six indices, followed by GPT-6 Astra (max) in Finance & Accounting, Strategy & Ops, Legal, and Engineering

➤ Open weights models are competitive across the Capability Indices. Kimi K3 (max) leads in Finance & Accounting (#8), Legal (#9), and Economics (#7), while DeepSeek V4.1 Flash (max) leads in Strategy & Ops (#7), and GLM-5.3 (max) leads in Healthcare & Medical (#6) and Engineering (#7)

来源:Artificial Analysis· x.com