跳到正文
ARC Prize· @arcprize · X·· 1 天前AI 评分44
AI 导读

ARC Prize 公布 DeepSeek-V4.1-Flash 在 ARC-AGI-1 上的结果:high 推理档得分 88.5%,低于 low 档的 90.5%,差异集中在 8 个任务上(high 在 5 个任务上低 4.5 分、3 个任务上高 2.5 分),且 high 多消耗 35% 输出 token 却未带来更稳定的正确答案。

正文

On ARC-AGI-1, high reasoning scores 88.5% vs low's 90.5%. Scores differed on 8 tasks: high scored worse on 5 (-4.5 points) and better on 3 (+2.5 points), accounting for the 2 percentage point drop. High used 35% more output tokens, but this didn't lead to consistently better answers.

On ARC-AGI-2, max reasoning has a slightly lower reported cost per task than high ($0.129 vs $0.133), but these averages divide recorded costs by all 120 tasks, including incomplete ones. API issues left 13 tasks incomplete with max reasoning and 8 tasks with high. Comparing the 106 complete tasks for both, max is actually 4.5% more expensive per task.

Full results: https://arcprize.org/results/deepseek-v4-1-flash

来源:ARC Prize · x.com