跳到正文
原文
The Decoder:AI News· Maximilian Schreiner·· 3 小时前同事件AI 评分77

Ataraxos 以85%有效胜率击败最强人类 Stratego 选手,训练成本不足 8000 美元

AI beats Stratego's greatest player, ending one of the last human strongholds in board games

AI 导读

研究人员开发的 AI 系统 Ataraxos 以 85% 的有效胜率(平局计半胜)击败了曾获 4 次世界冠军的史上最强 Stratego 选手 Niemeijer,并在 2025 世界锦标赛表演赛中取得 40 局 38 胜。

同一事件,精选展示《MIT 等机构发布 Ataraxos,以极低成本战胜顶级人类 Stratego 选手》

正文 · 原文

Niemeijer is considered the most decorated player in the game's history. He has won four world championships, 15 Dutch national titles, and two online world championships, and he spent more than 600 weeks at the top of the world rankings. George Franka, the only player who has competed in every world championship since 1997, calls Niemeijer "the best Stratego player of all time."

Stratego is an especially tough test for AI because both sides set up their 40 pieces face down. According to the paper, there are more than 10^33 possible setups. In games with imperfect information like this, the value of a move depends not only on how play continues afterward but also on what happened before and during the decision, first author Samuel Sokota wrote on X.

Both players set up their pieces in secret. A piece's rank is revealed only when it runs into an opposing piece. Whoever captures the enemy flag wins. | Image: Sokota, S., Vinitsky, E., Hu, H. et al.

According to Sokota, methods derived from poker AI did circumvent this problem, but only when there was little hidden information. In Texas Hold’em, there are only 1,326 possible starting hands, and according to the paper, the computational effort required by these methods increases with the amount of hidden information. Stratego was therefore one of the last major competitive board games in which humans remained superior to AI.

A university budget beat what DeepMind's millions couldn't

DeepMind had previously failed to push past the level of top human players in this very game. According to the paper, its DeepNash system trained on 1,024 TPU nodes for two to three months, as the corresponding DeepNash author recalls. The Ataraxos authors estimate that would cost $3 million to $4.5 million at 2025 prices.

Ataraxos, by contrast, needed one week on 16 Nvidia H100 GPUs, plus four more days on four GPUs for its so-called belief network. At 2025 prices, that comes to less than $8,000. The researchers put it at roughly 1/500th of the compute cost, 1/30th of the self-play games, and 1/100th of the training examples. They credit the custom GPU simulator they wrote, but also much higher sample efficiency. The team was limited to academic computing resources and picked up plenty of tricks over the years, Sokota writes.

There's no head-to-head comparison between the two systems. The researchers say they offered to build the necessary infrastructure, but DeepMind told them that wasn't possible because the DeepNash code no longer works. At the 2023 world championship, DeepNash won 19 of 28 games but lost to most of the top players, including Niemeijer.

Forcing the AI to mix things up keeps it from getting stuck

Ataraxos learns without any human data, purely by playing against itself. According to the paper, the key is how hard the system gets pushed to vary its play. An extra condition during training makes the AI spread out its setups and moves instead of locking into a fixed strategy early on. This is known as regularization. It matters in Stratego because strong play requires a lot of randomness, the authors say, and predictable players get exploited.

The researchers loosen this pressure as training goes on and adjust the size of the learning steps to match. Early on, Ataraxos varies its play heavily and learns in big jumps. Later, it plays more deliberately and only makes small corrections. This keeps the learning process from going in circles or turning chaotic, which tends to happen in games with imperfect information. The researchers compare regularization to an energy reserve. If the AI burns through it too fast, it makes quick early progress but then loses the ability to keep learning and often becomes easy to exploit.

Ataraxos also uses a belief network that predicts what the opponent's hidden pieces are. Before each move, the AI uses it to generate possible game states, plays out candidate moves, and runs an extra learning step just for that decision. The authors say earlier work skipped this kind of search because it was considered too hard with so much hidden information.

The AI won big despite a built-in disadvantage

The researchers spread the series against Niemeijer over three weeks to avoid fatigue and give him time to prepare. Niemeijer knew Ataraxos wouldn't adapt to his style. He received $1,000 for taking part, plus $100 per win and $50 per draw, so he had a financial reason to play his best.

The effective win rate of 85 percent, counting draws as half a win, is unprecedented at the highest level, the authors write. Three-time world championship runner-up Max Roelofs says margins at the top are razor-thin because the game forces players to take risks. The AI also played at a structural disadvantage. Niemeijer could adjust to it over many games, but it couldn't adjust to him. According to the paper, three-time world champion Vincent de Boer considers that a major handicap.

At an exhibition during the 2025 Stratego World Championship, Ataraxos also won 38 of 40 games against tournament players. Players describe its style as hard to read. The AI pulls bluffs that humans consider too risky, and when it's behind, it stalls so stubbornly that some players find it rude. It also seems almost eerily lucky, since its pieces always seem to be in exactly the right spot. The games against Niemeijer are available online.

The method works well beyond Stratego

The researchers used the same method to build other systems. In the Barrage Stratego variant, their AI won four 50-game series against three of the four top-ranked players, all of them two-time Barrage world champions. In the cooperative card game Hanabi, it set new records in every variant, and in the two-player version it needed two orders of magnitude less compute than the previous leader. In the Chinese card game Dou dizhu, it beat PerfectDou and DouZero, the previous benchmark systems.

The authors draw a sweeping conclusion from these results. They argue that large amounts of hidden information are no longer an obstacle for reinforcement learning and search. That would put practical AI within reach for many strategic decision problems, as long as fast and accurate simulators can be built for them. The paper points to financial markets, military conflicts, and negotiations as examples of situations with imperfect information.

The researchers acknowledge one limitation. Because the search only mimics a single learning step, its performance can't keep improving just by adding more compute time. A more sophisticated search method could change that. The team has made the code for Ataraxos publicly available.

来源:The Decoder:AI News · the-decoder.com