Chess and poker demand different kinds of reasoning. Chess players can see the board; poker players must decide what to do without seeing opponents’ cards. DeepMind’s Student of Games (SoG) is designed to handle both kinds of challenge with a single learning system.
Why one system for different games is difficult
Games with perfect information reveal the full state of play, such as where every piece sits on a board. AlphaZero specializes in this setting and can play chess and Go at a superhuman level. Its approach uses the rules of a game and repeated self-play to improve decisions.
Games with imperfect information hide some of what matters. In poker, for example, players do not know the other players’ cards. Success can depend on concealing intentions, so searching through visible positions alone is not enough. DeepStack beat human poker professionals in 2016, and Facebook demonstrated a poker AI that beat five players simultaneously in a tournament in mid-2019.
These systems were specialists: AlphaZero did not play poker, and DeepStack did not play chess. SoG aims to bridge that gap by combining guided search, self-play learning and game-theoretic reasoning.
How Student of Games learns
Rather than rely on the same search method used by AlphaZero, SoG begins with a simple decision tree of possible strategies and plays games against itself. After a game, it considers how choosing differently in each situation might have changed the result.
This process is called growing-tree counterfactual regret minimization, or GT-CFR. As training continues, the decision tree grows. In plain terms, the system uses imagined alternative choices to refine its strategies over time, including in games where hidden information makes opponents’ possible actions part of the problem.
The researchers describe SoG in a paper published in Science as the “first algorithm to achieve strong empirical performance in large perfect and imperfect information games — an important step towards truly general algorithms for arbitrary environments.” That is a claim about progress toward a broader kind of game-playing algorithm, not evidence that SoG has mastered every game.
Where it performs well—and where it falls short
DeepMind trained SoG to play chess, Go, poker and Scotland Yard, then compared it with bots including AlphaZero, GnuGo, Stockfish and Slumbot. SoG won the most games in poker and Scotland Yard.
Chess and Go show the limits of its generality so far. SoG lost 99.5% of its games against AlphaZero in those two games. DeepMind says it nevertheless plays at a very high amateur level. Its ability to participate across game types is notable, but that breadth does not mean it matches specialist systems in every contest.
The comparison highlights the trade-off the work is exploring: a shared approach can cover settings with different information rules, while a dedicated system can remain much stronger in its specialty. SoG’s results suggest that techniques for visible and hidden information can be brought together, while also showing how much performance can vary by game.
A step toward broader learning systems
DeepMind’s work on board and video games is fundamental research that could have applications in other economically attractive areas of AI. The article does not specify those applications, but the underlying question is whether methods that learn to plan under different rules can be useful beyond games.
The researchers say further improvements may be possible. They also want to find out whether similar performance can be achieved with significantly fewer computing resources—a practical question for any approach that depends on training through repeated play.
An earlier version of the paper appeared on arXiv in 2021, when the system was called Player of Games. The project’s central ambition remains clear: build an algorithm that can adapt across games where all information is visible and games where important information is hidden.