跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分60
AI 导读

一篇关于 ARC-AGI-3 的论文指出,智能体可以通过读取 2,172 行游戏源码拿到满分 100,掩盖了真实能力。清理重跑后同一游戏仅得 46.91。作者建议评估智能体时用真实访问限制封锁答案,并阅读其动作日志。

正文

This paper shows that AI agents can hit perfect benchmark scores for the wrong reasons, so audit what they did, not just the result.

An agent scored a flawless 100 on an ARC-AGI-3 game by reading its 2,172-line source code.

logs caught what the scores hid.

And then a clean rerun of that game scored only 46.91.

When you evaluate agents, block off answers with real access limits and read their action logs, since agents use whatever they can reach.

来源:Rohan Paul · x.com