Kaggle Game Arena:在竞技游戏中评估 LLM 的策略能力
Game Arena: Strategic LLM Evaluation in Competitive Environments
Kaggle 推出 Game Arena,一个通过竞技游戏评估 LLM 的开放平台,让模型在结构化环境中正面对战,游戏强度随模型进化自然提升,避免性能饱和。首批三个试点环境为国际象棋、扑克和狼人杀,覆盖完全信息、不完全信息与多人博弈场景,用于系统研究模型的策略规划、适应能力与不确定性下的鲁棒性。平台提供完整竞赛结果与评估指标,并保证可复现性与透明度。
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org