Search arXivSearch

arXiv · 2404.06516

Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games

Abstract

In this work, we study potential games and Markov potential games under stochastic cost and bandit feedback. We propose a variant of the Frank-Wolfe algorithm with sufficient exploration and recursive gradient estimation, which provably converges to the Nash equilibrium while attaining sublinear regret for each individual player. Our algorithm simultaneously achieves a Nash regret and a regret bound of $O(T^{4/5})$ for potential games, which matches the best available result, without using additional projection steps. Through carefully balancing the reuse of past samples and exploration of new samples, we then extend the results to Markov potential games and improve the best available Nash regret from $O(T^{5/6})$ to $O(T^{4/5})$. Moreover, our algorithm requires no knowledge of the game, such as the distribution mismatch coefficient, which provides more flexibility in its practical implementation. Experimental results corroborate our theoretical findings and underscore the practical effectiveness of our method.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jing Dong, Baoxiang Wang, Yaoliang Yu. 2024-04-04. Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games. https://arxiv.org/abs/2404.06516

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Information Design with Heterogeneous Beliefs in Bayesian Congestion Games

In many engineered systems, agents make decisions under incomplete information, creating opportunities for a planner to influence decentralized behavior through signaling. We study how such signaling can be designed in parallel-network, affine latency congestion games when users may not interpret recommendations using the same beliefs assumed by the planner. To do so, we consider Bayesian congestion games with private recommendations and formulate a robust information design problem in which obedience must hold uniformly over a neighborhood of a nominal prior. This addresses the previously uncharacterized issue of whether obedience itself remains reliable under belief heterogeneity, rather than only under the single prior used at the design stage. We characterize policy-level robustness radii, identify regimes in which the robust obedience region remains nonempty, and analyze the resulting robustness--performance tradeoff through a robust value function whose optimal cost is monotone in the robustness requirement and whose local sensitivity is governed by the active obedience constraints.

cs.GT

Core stability recognition for minimum-cost spanning tree games: Parameterized perspective

Minimum-cost spanning tree game (MSTG) is a cooperative game played on an undirected edge-weighted graph $(G,w)$ representing the network, where each vertex corresponds to a player and each edge has an associated cost~$w$. A distinguished vertex $s \in V(G)$ represents the supply or source. For any coalition of players $S$, the characteristic cost function $c(S)$ is defined as the minimum cost of a spanning tree with respect to $w$, connecting exactly the vertices in $S \cup \{s\}$. In this paper we study the computational complexity of deciding core membership for MSTG. In general, deciding whether a given allocation is in the core is \textsf{coNP}-hard~(Faigle et al.,International Journal of Game Theory,1997). We study the core recognition problem under the name {\sc MSTG Core Non-Membership}. We extend the hardness to graphs which are very close to being planar. On the positive side, we present several algorithmic results within the framework of parameterized complexity. We show that {\sc MSTG Core Non-Membership} is fixed-parameter tractable when parameterized by the support size of the allocation. Turning into structural parameters of graphs, we show that the problem admits an FPT algorithm parameterized by treewidth and signed neighborhood diversity. Last but not least, we investigate kernelization. While in general graphs, under standard complexity-theoretical assumptions, {\sc MSTG Core Non-Membership} does not admit a polynomial kernel parameterized by the vertex cover number, we design a cubic kernel in planar graphs. Furthermore, in general graphs, we obtain quadratic kernel for signed neighborhood diversity and linear kernel for the parameter feedback edge number.

cs.GT

Condorcet-type properties of the linear ordering problem with ties

The Kemeny rule aggregates multiple strict rankings into a single strict ranking that minimizes the sum of its distances from the input rankings. The resulting optimization problem, called the Kemeny problem (\texttt{KP}), is a special case of the linear ordering problem (\texttt{LOP}). The Kemeny rule satisfies several desirable properties in social choice theory, including the extended Condorcet criterion (\texttt{XCC}). Ando et al. strengthened this result by introducing the strong Condorcet criterion (\texttt{SCC}) and showing that it holds for every optimal solution to an arbitrary \texttt{LOP} instance. Yoo and Escobedo extended the Kemeny rule to rankings with ties and showed that the resulting rule satisfies the non-strict extended Condorcet criterion (\texttt{NXCC}). This criterion gives a condition under which one candidate must be ranked strictly above another in every optimal solution. In this paper, we introduce the non-strict strong Condorcet criterion (\texttt{NSCC}), a counterpart of the \texttt{SCC} for rankings with ties, and show that it holds for every optimal solution to an arbitrary instance of the linear ordering problem with ties (\texttt{LOPT}). We also establish a complementary structural property that gives conditions under which two candidates must be tied in every optimal solution to an arbitrary \texttt{LOPT} instance.

cs.GT