arXiv · 2609.16751
Constant Swap Regret in General-Sum Games via Two-Scale Higher-Order Optimism
Abstract
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret in self-play, independent of the horizon $T$. With $n$ players and at most $m$ actions each, every player's individual swap regret is $O(\sqrt n\,m\log m\log^{5/2}(nm))$ at every finite horizon. The dynamics use the classical Blum-Mansour framework with optimism. Each player predicts the deviation gains, uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. Our new ingredients include a tailored row normalization map and a two-scale higher-order predictor. An adversarially robust variant, obtained through a generic common-prefix switching wrapper, preserves the self-play bound up to a universal constant and guarantees individual swap regret at most $7\sqrt{mT\log m}$ in the adversarial setting.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tung Mai. 2026-09-19. Constant Swap Regret in General-Sum Games via Two-Scale Higher-Order Optimism. https://arxiv.org/abs/2609.16751
Cite the original work for its findings. Save a collection to share your selection of sources.