Search arXivSearch

arXiv · 2210.03852

Stackelberg POMDP: A Reinforcement Learning Approach for Economic Design

Abstract

We introduce a reinforcement learning framework for economic design where the interaction between the environment designer and the participants is modeled as a Stackelberg game. In this game, the designer (leader) sets up the rules of the economic system, while the participants (followers) respond strategically. We integrate algorithms for determining followers' response strategies into the leader's learning environment, providing a formulation of the leader's learning problem as a POMDP that we call the Stackelberg POMDP. We prove that the optimal leader's strategy in the Stackelberg game is the optimal policy in our Stackelberg POMDP under a limited set of possible policies, establishing a connection between solving POMDPs and Stackelberg games. We solve our POMDP under a limited set of policy options via the centralized training with decentralized execution framework. For the specific case of followers that are modeled as no-regret learners, we solve an array of increasingly complex settings, including problems of indirect mechanism design where there is turn-taking and limited communication by agents. We demonstrate the effectiveness of our training framework through ablation studies. We also give convergence results for no-regret learners to a Bayesian version of a coarse-correlated equilibrium, extending known results to correlated types.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gianluca Brero, Alon Eden, Darshan Chakrabarti, Matthias Gerstgrasser, Amy Greenwald, Vincent Li, David C. Parkes. 2024-07-19. Stackelberg POMDP: A Reinforcement Learning Approach for Economic Design. https://arxiv.org/abs/2210.03852

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Control and Bribery in Stable Marriage and Stable Roommates: A Complete Complexity Landscape

We study control and bribery problems for stable matchings: A central authority (the controller, resp. briber) may add agents, delete agents, delete acceptable pairs, swap two adjacent agents in some agent's preference list, or arbitrarily reorder some agent's preference list, in an instance of Stable Marriage or Stable Roommates. We extend previous work on control and bribery in stable matchings by Boehmer et al. [8]. We consider goals capturing individual and pair inclusion, stability, and uniqueness requirements: Matching a designated agent (MA), matching a designated pair (MP), realizing a stable matching consistent with a given matching (MS), making a given matching the unique stable matching (USM), or guaranteeing that a stable (resp. perfect and stable) matching exists ($\exists$SM/$\exists$PSM). We provide a unified complexity map for all non-trivial action-goal combinations in both settings, consolidating known results and extending the study to the roommates model, where stable matchings need not exist.

cs.GT

How a Cooperative-Override Circuit Suppresses Nash Play in Large Language Models

On the named Prisoner's Dilemma under direct prompting, three larger instruction-tuned models, Llama-3-70B, Qwen2.5-32B, and Qwen2.5-72B, lock at full cooperation, the metric's maximum distance from Nash with zero variance across replicates, while Llama-3-8B plays near-Nash. Opening the models, a logit-lens analysis finds a distributed cooperative override. Intermediate readouts lean toward the Nash action through roughly three quarters of network depth before a late surge toward cooperation, and the final layer settles the contest. The size of that final correction, not the surge, rank-matches chain-of-thought behavior across scale and two architectures. In the 8B the override is a single causally controllable direction in the residual stream; steering it dials the decision, and clamping its component at one position of one layer moves the choice strictly monotonically, Spearman rho = 1.000, with generation fluent. The circuit is lexical. It survives name removal and payoff rescaling but disengages when Cooperate and Defect are replaced with neutral labels, and on 48 payoff-random games with neutral surfaces no model locks cooperative on any dilemma or shows general equilibrium competence. In mixed-model populations a single Nash-playing agent collapses cooperation contagiously. What suppresses Nash play in large language models is a word-triggered circuit rather than missing competence, and it can be measured, bounded, and controlled.

cs.GT

Auction Design with ROI-Constrained Bidders: Truthfulness and Revenue Maximization

The return-on-investment (ROI) constraint is central to many auctions, particularly in online advertising, where a bidder is unwilling to pay more than a fixed fraction of the value obtained. We study truthful and revenue-maximizing auctions for ROI-constrained bidders. We first characterize truthful auctions when both valuations and ROI constraints are private, showing that the allocation rule uniquely determines the payment rule. Building on this characterization, for multiple bidders we introduce $σ$-increment mechanisms that resemble Myerson's optimal mechanism~\cite{journals/mor/Myerson81}; as $σ$ vanishes, these mechanisms become asymptotically optimal among deterministic truthful mechanisms, and their revenue approaches at least a $1/\bar r$ fraction of the optimal expected revenue over all truthful mechanisms, where $\bar r$ is the largest possible ROI constraint. In the single-bidder setting, we prove that every truthful auction can be replaced by a convex pricing function with weakly higher payments for every type, and we derive the optimal pricing functions when either the valuation or the ROI constraint is public.

cs.GT