Agent Arena 分析 29 模型 20840 条编码反馈

Arena.ai · @arena · X·2026-09-21 23:38·2小时前
AI 导读

Arena 分析 Agent Arena: Code 中 29 个模型的 20840 条 trace 后发现,前沿模型获得的正向反馈明显更多,新模型整体反馈更偏正面。

Arena.ai@arena
44AI 编辑部评分,满分 100

Agent Arena 分析 29 模型 20840 条编码反馈

2026-09-21 23:38· 2小时前
AI 导读

Arena 分析 Agent Arena: Code 中 29 个模型的 20840 条 trace 后发现,前沿模型获得的正向反馈明显更多,新模型整体反馈更偏正面。

What can user praise and complaints tell us about coding agents themselves?

We analyzed 20,840 traces from Agent Arena: Code across 29 models, and found that direct feedback unlocks novel opportunities in tracking the frontier.

Some topline findings: • Today’s frontier models receive markedly more positive feedback. Overall, newer models across labs tend to show a more positive feedback balance. • Broken code is still the biggest driver of complaints, sloppy behavior comes second. 69.1% of complaints point to code not working. The next themes are incomplete output (27.1%), weak finish and usability (27.1%), and ignored instructions (20.3%). • Models share common weaknesses, but differ in how they disappoint. Fable 5.1 attracts fewer “slop” and design complaints than Astra (3.9% versus 5.7% of sampled traces), and fewer complaints about being “flaky” (3.1% versus 5.0%).

Read more by diving into the full article from @DawidGalarowicz below.

Arena.aihttps://x.com/i/article/2101066086867492864

来源:Arena.ai· x.com