Artificial Analysis Coding Agent Index 新增安全拒答报告

Artificial Analysis · @ArtificialAnlys · X·2026-09-19 07:26·59分钟前
AI 导读

Artificial Analysis 在 Coding Agent Index v1.5 中新增安全拒答报告,用于解释模型行为与得分差异。Claude Fable 5.1 在 Claude Code 和 Devin Fusion 中回退率最高,回退尝试分别占指数权重的 8.8% 和 7.1%,结果因此包含回退模型的表现。

Artificial Analysis@ArtificialAnlys
41AI 编辑部评分,满分 100

Artificial Analysis Coding Agent Index 新增安全拒答报告

2026-09-19 07:26· 59分钟前
AI 导读

Artificial Analysis 在 Coding Agent Index v1.5 中新增安全拒答报告,用于解释模型行为与得分差异。Claude Fable 5.1 在 Claude Code 和 Devin Fusion 中回退率最高,回退尝试分别占指数权重的 8.8% 和 7.1%,结果因此包含回退模型的表现。

Safety refusal reporting is now available in the Artificial Analysis Coding Agent Index

In our latest Coding Agent Index v1.5, we’ve introduced safety refusal reporting to help explain model behavior and score differences. A safety refusal occurs when a provider or model declines to start or continue a task on safety grounds.

An agent may fall back to another model to continue, or stop the attempt with a block.

Claude Fable 5.1 had the highest fallback rates in both Claude Code and Devin Fusion, with fallback attempts accounting for 8.8% and 7.1% of the Index's weight, respectively; these results therefore include the fallback models' performance.

Refusal variability, harness context buildup, effort settings, and retry strategies can all affect the observed rates.

来源:Artificial Analysis· x.com