What can user praise and complaints tell us about coding agents themselves?
We analyzed 20,840 traces from Agent Arena: Code across 29 models, and found that direct feedback unlocks novel opportunities in tracking the frontier.
Some topline findings: • Today’s frontier models receive markedly more positive feedback. Overall, newer models across labs tend to show a more positive feedback balance. • Broken code is still the biggest driver of complaints, sloppy behavior comes second. 69.1% of complaints point to code not working. The next themes are incomplete output (27.1%), weak finish and usability (27.1%), and ignored instructions (20.3%). • Models share common weaknesses, but differ in how they disappoint. Fable 5.1 attracts fewer “slop” and design complaints than Astra (3.9% versus 5.7% of sampled traces), and fewer complaints about being “flaky” (3.1% versus 5.0%).
Read more by diving into the full article from @DawidGalarowicz below.