We just rolled out CursorBench 4.0!
It includes new tasks for how well models follow instructions, work on challenging projects over time, and is more difficult than before (so all models score lower).
我们刚刚推出了 CursorBench 4.0! 它新增了评估模型指令遵循能力、长期处理高难度项目的任务,并且比之前更难(所以所有模型的得分都更低了)。
我们刚刚推出了 CursorBench 4.0! 它新增了评估模型指令遵循能力、长期处理高难度项目的任务,并且比之前更难(所以所有模型的得分都更低了)。
We just rolled out CursorBench 4.0!
It includes new tasks for how well models follow instructions, work on challenging projects over time, and is more difficult than before (so all models score lower).
来源:Lee Robinson· x.com