SlopCodeBench 完整跑分:GLM 5.3、Sol 5.6 与 Astra 对比

dex · @dexhorthy · X·2026-09-13 04:02·46分钟前
AI 导读

Dex Horthy 完成了 SlopCodeBench 对 GLM 5.3、Sol 5.6 和 Astra 的完整基准测试,首次跑完全部场景而非此前的小子集。Astra 得分比 gpt-5.5 高几分,但差距不如预期;Sol 得分低于 gpt-5.5 令他意外。他猜测新模型在更高思考模式下更容易失控,自己日常多用 Sol 的 medium 或 low effort。

dex@dexhorthy
30AI 编辑部评分,满分 100

SlopCodeBench 完整跑分:GLM 5.3、Sol 5.6 与 Astra 对比

2026-09-13 04:02· 46分钟前
AI 导读

Dex Horthy 完成了 SlopCodeBench 对 GLM 5.3、Sol 5.6 和 Astra 的完整基准测试,首次跑完全部场景而非此前的小子集。Astra 得分比 gpt-5.5 高几分,但差距不如预期;Sol 得分低于 gpt-5.5 令他意外。他猜测新模型在更高思考模式下更容易失控,自己日常多用 Sol 的 medium 或 low effort。

finally finished a full SlopCodeBench run against GLM 5.3, Sol 5.6, and Astra (Fable 5.1 results coming soon)

This is different from all our previous research on SlopCodeBench (from @GOrlanski @ UW) in that we ran the full benchmark, every single scenario here. Previous runs did a small subset of the challenges.

Asterisks: • this was run over ~1 week and had to be resumed a few times due to various provider outages • I still think a more scientific approach here would be to do what's common with other benchmarks, which is run multiple evaluations and aggregate the scores

Looks like Astra scores a few points higher than gpt-5.5 here. Not as many as I'd think. I'm surprised Sol got lower than gpt 5.5 because in my experience I like working with Sol a bit more.

My vibes-best guess is that the newer models are likely to go more off the rails on higher thinking modes (e.g. I almost always use Sol in medium or low effort for most work on @humanlayer_dev)

来源:dex· x.com