arXiv · 2609.36787
Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions
Abstract
Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ondřej Kubíček, Viliam Lisý, Tuomas Sandholm. 2026-09-29. Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions. https://arxiv.org/abs/2609.36787
Cite the original work for its findings. Save a collection to share your selection of sources.