Wednesday, October 7, 2026
English edition

Development

With most information hidden, the game Stratego

October 1, 2026 Development Source: Ars Technica

With most information hidden, the game Stratego

Share this article

Just like DeepNash, Ataraxos learned by playing against itself—163 million games in total. In these self-play sessions, moves that led to wins were reinforced and played more often in future matches, while moves that led to losses were played less, which was the same simple training idea. The difference was in how much Ataraxos adjusted after each game, because hidden information tends to send self-play learning algorithms around in circles. The team addressed this by making big, bold changes in strategy early in training and small, careful ones later. The even bigger innovation was something DeepNash never had: thinking ahead before each move. AIs like AlphaGo refine their general strategy with a search just before acting. DeepMind couldn’t make that work in Stratego because the search space was too large, leaving it an open question whether it was worth trying. “This is one of the things that we did figure out how to do,” Farina said. The solution was a second neural network, a belief model, trained to guess the opponent’s hidden pieces based on how they had been moving. This way, instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on how they turned out. The name Ataraxos comes from the ancient Greek word for a state of calm. “It means somebody that’s calm and unbothered,” Farina explained. He suggests the structure of the AI and its lack of human emotions ensure it doesn’t react impulsively, “even in situations where a human would be losing their mind.” While the human might try big gambles to come back from a significant deficit, Ataraxos would work its way back into the game slowly and methodically. But Ataraxos’s best trick was arguably its price tag. DeepNash was trained for two to three months on 1,024 of Google’s specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos, in contrast, needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model. Farina and lead author Samuel Sokota achieved this efficiency by writing a simulator that runs millions of moves per second on graphics cards. “At the scale that we are in academia, we don’t really have access to an entire field of GPUs,” Farina said. The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger. The Ataraxos architecture also worked in learning other games. The same approach beat three world champions at Barrage Stratego, a faster eight-piece variant of Stratego, mastered the cooperative card game Hanabi, and beat the best bots at the Chinese card game dou dizhu. But the team has its sights set on scenarios far more complex than board or card games. Board games have fixed rules and clear winners, while real-world problems like negotiations, financial markets, or military conflicts usually don’t. But the Ataraxos team argues the gap is smaller than it looks, since tackling any real problem starts with building a simplified model of it. “War gaming is a common thing that people do,” Vinitsky said. “You can use the techniques that were derived here to play it forward and see how a strong opponent might respond to what you do.” Farina and his colleagues are now interested in making their AI more understandable, since Ataraxos, in its current state, can’t explain why it makes the moves it makes. “We work on machines that produce strong but also interpretable and explainable strategies. I think we’re not quite there yet,” Farina said. Nature, 2026. DOI: 10.1038/s41586-026-11036-y