Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations
The Capability Indices map tasks from O*NET occupations to benchmarks that represent them, weighting each benchmark by how often its capability appears across the work. We cover six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics
Key changes: ➤ Finance & Accounting: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Finance), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work
➤ Strategy & Ops: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work
➤ Legal: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work
➤ Healthcare & Medical: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Long-Context Reasoning added as a new capability sourced from MLCR-AA, Agentic Customer Interaction removed from capabilities, and AA-Briefcase added to Agentic Knowledge Work
➤ Engineering: Terminal-Bench updated to v4.0 in Agentic Terminal Use, GPQA Diamond removed from Reasoning, and AA-Briefcase added to Agentic Knowledge Work
➤ Economics: AA-Briefcase added to Agentic Knowledge Work
Key results: ➤ Claude Fable 5.1 (max) leads in all six indices, followed by GPT-6 Astra (max) in Finance & Accounting, Strategy & Ops, Legal, and Engineering
➤ Open weights models are competitive across the Capability Indices. Kimi K3 (max) leads in Finance & Accounting (#8), Legal (#9), and Economics (#7), while DeepSeek V4.1 Flash (max) leads in Strategy & Ops (#7), and GLM-5.3 (max) leads in Healthcare & Medical (#6) and Engineering (#7)