Search arXiv⌕ Search

arXiv · 2204.04930

A Unified Perspective on Deep Equilibrium Finding

Abstract

Extensive-form games provide a versatile framework for modeling interactions of multiple agents subjected to imperfect observations and stochastic events. In recent years, two paradigms, policy space response oracles (PSRO) and counterfactual regret minimization (CFR), showed that extensive-form games may indeed be solved efficiently. Both of them are capable of leveraging deep neural networks to tackle the scalability issues inherent to extensive-form games and we refer to them as deep equilibrium-finding algorithms. Even though PSRO and CFR share some similarities, they are often regarded as distinct and the answer to the question of which is superior to the other remains ambiguous. Instead of answering this question directly, in this work we propose a unified perspective on deep equilibrium finding that generalizes both PSRO and CFR. Our four main contributions include: i) a novel response oracle (RO) which computes Q values as well as reaching probability values and baseline values; ii) two transform modules -- a pre-transform and a post-transform -- represented by neural networks transforming the outputs of RO to a latent additive space (LAS), and then the LAS to action probabilities for execution; iii) two average oracles -- local average oracle (LAO) and global average oracle (GAO) -- where LAO operates on LAS and GAO is used for evaluation only; and iv) a novel method inspired by fictitious play that optimizes the transform modules and average oracles, and automatically selects the optimal combination of components of the two frameworks. Experiments on Leduc poker game demonstrate that our approach can outperform both frameworks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xinrun Wang, Jakub Cerny, Shuxin Li, Chang Yang, Zhuyun Yin, Hau Chan, Bo An. 2022-04-11. A Unified Perspective on Deep Equilibrium Finding. https://arxiv.org/abs/2204.04930

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Dynamic Welfare-Maximizing Pooled Testing

Pooled testing uses one test to certify several agents as healthy when the pooled result is negative. We study a budget-constrained welfare problem in which agents have heterogeneous utilities and independent prior probabilities of being healthy. Welfare is earned when an agent is certified healthy, and a dynamic policy may choose each pool after observing earlier test outcomes. We ask how much such adaptation can improve over a static allocation that fixes all pools in advance. Our main result proves that the optimal dynamic policy has value at most twice that of the optimal static overlapping allocation, for every population, test budget, and pool-size cap. The proof samples a path through the dynamic policy using an independent health profile, randomly thins the selected pools, and compares the resulting static allocation with the dynamic policy one agent at a time. A Boolean-cube argument proves the comparison on uniform subcubes, and a complementary-profile coupling lifts the result to arbitrary heterogeneous product priors. We also identify regimes in which adaptivity has no value, show that strict adaptive gains require re-pooling agents after positive tests, and give a three-agent instance in which adaptation is strictly beneficial. Exact-Joint Greedy obtains a $1/(e+1)$ fraction of optimal static overlapping welfare. The static non-overlapping Greedy algorithm of Finster et al. has the same guarantee, hence our factor-two theorem newly implies that each is a $2(e+1)$-approximation to the optimal dynamic policy. Finally, a reproducible exact small-instance study compares the dynamic and static benchmarks and greedy policies. The appendix records separate exploratory results for Gibbs-marginal and reinforcement-learning approaches at larger scales.

cs.GT↗

Induced Representations in Cooperative Games with Homogeneous Groups of Players

Oftentimes, the Shapley value, a measure of the contribution of a player to a game, becomes infeasible for games with many players. However, establishing symmetry allows for polynomial-time computation. To examine this reduction, we identify the spectrum of a homogeneous group game by using an induced representation from a Young subgroup. We prove that the depth of interaction of a two-group game is limited by the size of the minority group. Therefore, the algebraic structure of the game filters out a large space of irrelevant complexities. We then show that this filtration constrains any symmetric linear value to a specific subspace. This recovers the Shapley value uniquely for games consisting of exactly two homogeneous groups under standard axioms. Finally, we explore applications to the UN Security Council and complementary goods markets to illustrate the practical power of this approach.

cs.GT↗

From Bilateral Trade to Matching Markets: Sharp Gains from Trade

We study gains from trade in matching markets with independent private values and costs, Bayesian incentive compatibility, interim individual rationality, and no expected budget deficit. A second-best guarantee for finite bilateral trade extends without loss to matching markets with independent Borel priors, arbitrary downward-closed feasibility, and finite expected first-best gains. For bounded buyers with monotone hazard rates and arbitrary bounded sellers, we determine the exact worst-case ratio of second-best to first-best gains, approximately $0.72490721$. For binary buyers and sellers with at most $m$ types, we determine the exact ratio for every $m$, including $8/9$ when $m=2$ and a limit of $4/5$ as $m$ grows. Both families of bounds are tight already in bilateral trade.

cs.GT↗