跳到正文
ARC Prize· @arcprize · X·· 3 小时前AI 评分32
AI 导读

@SpaceXAI Grok 4.7 的推理 token 用量与其 ARC-AGI-2 公开得分相关:low 档平均每次测试对尝试用 10k token,得分 25%;而 medium 到 xhigh 档用 86k 到 120k token,得分 57.5% 到 60%。low 档更低的 token 用量可能有助于解释其更低的得分。

正文

@SpaceXAI Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public score: low averaged 10k tokens per test-pair attempt and scored 25%, versus 86k to 120k tokens and 57.5% to 60% at medium through xhigh. Low's lower token usage may help explain its lower score.

来源:ARC Prize · x.com