跳到正文
ARC Prize· @arcprize · X·· 3 小时前AI 评分40
AI 导读

Grok 4.7 在 ARC-AGI-2 semi-private 任务的中、高、xhigh 推理档位上平均消耗的 reasoning token 均多于 Grok 4.6,导致每任务成本更高。

正文

Grok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning, contributing to its higher cost per task. Per test-pair attempt, medium used 136% more reasoning tokens, high 125% more, and xhigh 173% more. Low reasoning used 27% fewer.

The chart below compares both models at xhigh on the 20 ARC-AGI-2 public tasks where Grok 4.7 used the most reasoning tokens per attempt, ordered by the token increase over Grok 4.6. Each circle shows the average across all recorded attempts for all test pairs in that task.

来源:ARC Prize · x.com