GPT-6.1 Sol 在 ARC-AGI-3 上以更高推理等级做出更优决策,完成关卡所需动作更少,从而降低总推理成本。以 sk48 为例,使用保留不透明推理状态并支持自动压缩的 provider adapter 时,max 推理仅用 463 个动作通关,而 low 推理用了 1,078 个。
Like GPT-6 Astra, GPT-6.1 Sol made better decisions at higher reasoning levels on ARC-AGI-3, requiring fewer actions to complete levels, which reduced the total inference cost.
For example, on sk48 using the provider adapter (which preserves opaque reasoning state between turns and enables auto compaction), GPT-6.1 Sol with max reasoning worked out how to reposition the colored blocks around an obstacle sooner, while low spent much longer revisiting blocked moves and testing controls that didn't help. That helped it complete the game in 463 actions versus 1,078 with low reasoning.
Watch the replay: https://arcprize.org/replay/c299de73-6d1c-45f5-be74-c3552c90af73
来源:ARC Prize · x.com