An AI system developed by researchers from MIT, Carnegie Mellon, NYU and Stanford has beaten top Stratego players, including the world’s strongest player, while using far fewer training examples than its predecessor. The result is a striking advance in hidden-information game-playing, with possible lessons for other strategic decisions where nobody gets the full picture.
MIT CSAIL Watch analysis
What happened
The researchers’ system, Ataraxos, combines self-play reinforcement learning with decision-time planning. It estimates the likely identities of an opponent’s concealed pieces, then refines its move as the game unfolds.
In results described by MIT News, Ataraxos beat the world’s strongest Stratego player 15-1-4 and recorded a 39-2 record against top players at the Stratego world championship. The researchers say it surpassed DeepNash, an earlier system, using less than one hundredth of its training examples and less than one thirtieth of its self-play games.
The team also adapted the system to Barrage Stratego, Hanabi and Dou dizhu, where it achieved superhuman performance. Read the MIT research report.
Why it matters
Stratego has more than 10⁶⁶ possible piece configurations, and players must reason about what their opponent knows as well as what they might do. That makes it a demanding test of decision-making under uncertainty, not simply a matter of searching a large game tree.
The researchers point to possible applications in areas such as business negotiations and cybersecurity. Those are prospective uses, not demonstrated real-world results: winning a game does not by itself make an AI a trustworthy adviser in situations with people, shifting incentives and consequences beyond the scoreboard.
Our read
The notable achievement is the combination of stronger play with a much smaller training bill. That is more interesting than a machine winning another game, although the games still provide the cleanest evidence for what Ataraxos can do.
The team says it wants to develop ways to make the system’s decisions interpretable. That is the right next move: before anyone asks for strategic advice, they should be able to see why the machine thinks a bluff is a bluff.
What to watch
- Whether the researchers publish further details on Ataraxos’s methods and evaluation.
- Whether its performance holds up in more games and less controlled decision-making settings.
- How the team develops ways to audit and explain the system’s recommendations.
Discussion spark: Would you trust an AI with a strong record in hidden-information games to advise on real negotiations or cybersecurity, or are those settings too different for the win record to carry weight?
Sources and evidence
- This game-playing AI is the new champ at Stratego (30 September 2026, 15:00 UTC)
Independent WittyWires tracker for public updates about MIT CSAIL. Not affiliated with or endorsed by MIT CSAIL; this is not an official account.