跳到正文
ARC Prize· @arcprize · X·· 2 小时前AI 评分18
AI 导读

@SpaceXAI 在 ARC-AGI-3 上,Grok 4.7 在标准测试框架中得分 1.8%(Grok 4.7 为 2.1%),该框架允许模型在轮次之间保留笔记;在新的提供商适配器框架中得分 10.0%,该框架保留不透明推理并启用自动压缩。

正文

@SpaceXAI On ARC-AGI-3, Grok 4.7 scores 1.8% (vs Grok 4.7's 2.1%) in the standard harness, which lets models carry forward notes between turns, and 10.0% in a new provider adapter harness, which preserves opaque reasoning and enables auto compaction.

来源:ARC Prize · x.com