arXiv · 2303.00177
Finite-sample Guarantees for Nash Q-learning with Linear Function Approximation
Abstract
Nash Q-learning may be considered one of the first and most known algorithms in multi-agent reinforcement learning (MARL) for learning policies that constitute a Nash equilibrium of an underlying general-sum Markov game. Its original proof provided asymptotic guarantees and was for the tabular case. Recently, finite-sample guarantees have been provided using more modern RL techniques for the tabular case. Our work analyzes Nash Q-learning using linear function approximation -- a representation regime introduced when the state space is large or continuous -- and provides finite-sample guarantees that indicate its sample efficiency. We find that the obtained performance nearly matches an existing efficient result for single-agent RL under the same representation and has a polynomial gap when compared to the best-known result for the tabular case.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pedro Cisneros-Velarde, Sanmi Koyejo. 2023-03-01. Finite-sample Guarantees for Nash Q-learning with Linear Function Approximation. https://arxiv.org/abs/2303.00177
Cite the original work for its findings. Save a collection to share your selection of sources.