MIT researchers challenge a long-held assumption in game theory: that specialized game-theoretic algorithms outperform general-purpose policy gradient methods in two-player imperfect-information games. Their study, presented at ICLR 2025, shows that neural networks trained with policy gradient methods actually achieve better exploitability scores than those trained with game-theoretic algorithms. A key contribution is a new benchmark for fairly evaluating algorithms on imperfect-information games with up to 30 billion states — runnable on a laptop via a single line added to OpenSpiel. The findings have implications beyond recreational games, extending to military, trading, and negotiation scenarios.
45 Impressions