跳到正文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分46
AI 导读

清华论文构建 AAArena,用清华年度机器人竞赛的 12 款游戏和 1920 个存档人类程序作为对手,测试编码智能体能否在模型权重不变的情况下通过读规则、选对手、研究回放来改写游戏 bot。在全部 3 款测试游戏中,详细回放反馈均优于仅胜负信号;有回放时 Pacman bot 排名第 1,无回放仅第 11。但比赛预算增至 3 倍,4 个卡住的 bot 仍无一升至第 1。

正文

New Tsinghua paper finds that AI agents improving game bots from match replays can top human leaderboards, but mostly stall on games with complex rules.

Getting AI to learn a winning game strategy from a limited number of matches is still hard, especially against changing rivals.

They built AAArena from 12 games in Tsinghua's yearly bot-building contest, with 1,920 archived human programs as rivals. A coding agent, with its model weights unchanged, reads the rules, picks opponents, studies replays, and rewrites its bot within a match budget.

Detailed replays beat win/loss-only feedback in all 3 games tested. With replays, a Pacman bot reached rank 1, versus rank 11 without them.

Tripling the match budget did not push any of 4 stuck bots to rank 1.

– arxiv. org/abs/2610.12341

Title: "Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition"

来源:Rohan Paul · x.com