finally finished a full SlopCodeBench run against GLM 5.3, Sol 5.6, and Astra (Fable 5.1 results coming soon)
This is different from all our previous research on SlopCodeBench (from @GOrlanski @ UW) in that we ran the full benchmark, every single scenario here. Previous runs did a small subset of the challenges.
Asterisks: • this was run over ~1 week and had to be resumed a few times due to various provider outages • I still think a more scientific approach here would be to do what's common with other benchmarks, which is run multiple evaluations and aggregate the scores
Looks like Astra scores a few points higher than gpt-5.5 here. Not as many as I'd think. I'm surprised Sol got lower than gpt 5.5 because in my experience I like working with Sol a bit more.
My vibes-best guess is that the newer models are likely to go more off the rails on higher thinking modes (e.g. I almost always use Sol in medium or low effort for most work on @humanlayer_dev)