PublicationsAlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal PlayVlad Murgoci, Matthijs Spaan, and Yaniv Oren. AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play. arXiv:2605.09150, 2026. DownloadAbstractPoker is an imperfect information game that has served as a long-standing benchmark for decision-making under uncertainty. To maximize utility beyond the Nash equilibrium, an agent can deviate from Nash-equilibrium policies to exploit suboptimal play. We introduce AlphaExploitem, which extends the competitive RL poker agent AlphaHoldem by using a hierarchical transformer encoder that enables reasoning over previously played hands and modifying the training procedure with the inclusion of a diverse pool of exploitable opponents to facilitate learning to exploit. We train and evaluate AlphaExploitem on two standard benchmarks for imperfect-information games. Empirically, AlphaExploitem successfully exploits weak play by both in- and out-of-distribution opponents, without losing performance against NE opponents. BibTeX Entry@Misc{Murgoci26arxiv,
author = {Vlad Murgoci and Matthijs Spaan and Yaniv Oren},
title = {{AlphaExploitem}: Going Beyond the {N}ash
Equilibrium in Poker by Learning to Exploit
Suboptimal Play},
howpublished = {arXiv:2605.09150},
year = 2026
}
Note: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted without the explicit permission of the copyright holder. Generated by bib2html.pl (written by Patrick Riley) on Fri Aug 28, 2026 12:56:06 UTC |