Ataraxos AI Beats Top Stratego Player Using Belief Network on a Budget

Ataraxos defeats the best human Stratego player

The AI system named Ataraxos beat the world‑champion Stratego player Pim Niemeijer 15 games to 1 (with four draws), proving that imperfect‑information board games can be mastered without massive compute budgets.

Why hidden information makes Stratego hard for AI

Stratego hides the identity of each piece until a battle occurs, creating a "massive amount of hidden information" that unfolds over thousands of moves. With 40 pieces that can be arranged in more than a decillion configurations, the game’s state space is far larger than poker’s 1,326 possible hands. This depth of uncertainty, combined with long game lengths (often >2,000 moves) and bluffing dynamics, prevented earlier AI efforts such as DeepMind’s DeepNash from achieving reliable superhuman play.

Core technical breakthrough: a belief‑model neural network

Ataraxos adds a second neural network that predicts the opponent’s hidden piece identities from movement patterns. Rather than enumerating every possible board arrangement, the belief model samples plausible hidden states, runs a limited search over candidate moves in each sampled state, and selects the move with the best expected outcome. This approach enables forward search in a game where the full information tree is astronomically large.

"The solution was a second neural network, a belief model, trained to guess the opponent’s hidden pieces based on how they had been moving." – Gabriele Farina, MIT co‑author

Training efficiency and scale

  • Self‑play games: 163 million total, ~34 × fewer than DeepNash.
  • Compute: 16 GPUs for one week plus four GPUs for four days to train the belief model.
  • Cost: a few thousand US dollars, compared with DeepNash’s estimated $3–4.5 million on 1,024 Google‑custom chips.
  • Simulator: custom GPU‑accelerated engine runs millions of moves per second, allowing rapid reinforcement‑learning updates.

Gameplay characteristics of Ataraxos

  • Calm, methodical play – named after the Greek word ataraxos (tranquility), the AI avoids impulsive gambles and slowly recovers from low‑probability positions.
  • Effective bluffing – the belief model lets the AI make moves that would only be rational if the opponent were bluffing, and then follow through more reliably than humans.
  • Unconventional flag placement – the bot frequently hides its flag behind just two bombs in a corner, a setup rarely used by human experts, forcing opponents to adapt.

Match results against Pim Niemeijer

  • 20 online games over three weeks; Niemeijer earned $100 per win.
  • Final score: 15 wins for Ataraxos, 1 win for Niemeijer, 4 draws.
  • The single human win is attributed to luck; even perfect play can lose due to random initial piece arrangements.

Broader impact and extensions

The same architecture succeeded in:

  • Barrage Stratego – an eight‑piece fast variant, beating three world champions.
  • Hanabi – a cooperative card game where players cannot see their own cards.
  • Dou dizhu – a popular Chinese card game, where Ataraxos outperformed top bots.

The researchers suggest the techniques could transfer to real‑world domains such as war‑gaming, negotiations, or financial markets, where building a simplified model of hidden information is the first step.

Remaining challenges

  • Interpretability – Ataraxos cannot currently explain why it selects a particular move; the team is working on making strategies more transparent.
  • Generalization – applying the belief‑model search to games with dynamic rule sets or continuously expanding action spaces (e.g., Magic: The Gathering) remains an open problem.

Community reaction (Hacker News highlights)

  • Users noted the dramatic reduction in compute compared with DeepNash and praised the belief‑model insight as the critical innovation.

    "The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger. Imo, this is the critical piece and what makes the AI work at all." – @janalsncm

  • Several commenters recalled personal Stratego experiences and expressed surprise that the game had resisted AI longer than chess or Go.
  • Some raised questions about variant rules (e.g., "silent defense") and whether the approach could extend to other imperfect‑information games like bridge.

Future directions

The authors plan to improve explainability, explore larger‑scale hidden‑information domains, and release the Ataraxos codebase (available at https://ataraxosai.github.io/). Their results demonstrate that sophisticated belief modeling can close the gap between human intuition and machine calculation without prohibitive hardware costs.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Project
  • Dispatch