Search arXivSearch

arXiv · 2511.19828

Expectation-enforcing strategies for repeated games

Abstract

Originating in evolutionary game theory, the class of "zero-determinant" strategies enables a player to unilaterally enforce linear payoff relationships in simple repeated games. An upshot of this kind of payoff constraint is that it can shape the incentives for the opponent in a predetermined way. An example is when a player ensures that the agents get equal payoffs. While extensively studied in infinite-horizon games, extensions to discounted games, nonlinear payoff relationships, richer strategic environments, and behaviors with long memory remain incompletely understood. In this paper, we provide necessary and sufficient conditions for a player to enforce arbitrary payoff relationships (linear or nonlinear), in expectation, in discounted games. These conditions characterize precisely which payoff relationships are enforceable using strategies of arbitrary complexity. Our main result establishes that any such enforceable relationship can actually be implemented using a simple two-point reactive learning strategy, which conditions on the opponent's most recent action and the player's own previous mixed action, using information from only one round into the past. For additive payoff constraints, we show that enforcement is possible using even simpler (reactive) strategies that depend solely on the opponent's last move. In other words, this tractable class is universal within expectation-enforcing strategies. As examples, we apply these results to characterize extortionate, generous, equalizer, and fair strategies in the iterated prisoner's dilemma, asymmetric donation game, nonlinear donation game, and the hawk-dove game, identifying precisely when each class of strategy is enforceable and with what minimum discount factor.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nikos Dimou, Alex McAvoy. 2025-11-25. Expectation-enforcing strategies for repeated games. https://arxiv.org/abs/2511.19828

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Information Greenhouse: Optimal Persuasion for Medical Test-Avoiders

Patients often avoid medical tests because the information they provide, although medically useful, is psychologically painful. This paper studies optimal communication between a doctor and an information-avoidant patient who can refuse testing and treatment. I characterize when optimal communication creates an information greenhouse, a commitment to reward participation with comforting information about the untreated prognosis. When testing is voluntary and the patient is unwilling to be tested under extreme pessimism, an information greenhouse is optimal and takes the form of committed comfort, which provides reassuring information after the test. When the patient can reject the consultation at the outset and the patient's prior belief about the untreated prognosis is intermediate, an information greenhouse is optimal and takes the form of precautionary comfort, which provides reassuring information before the test. In all other cases in which the patient can be persuaded, warning-based policies that trigger pessimism prevail.

econ.TH

Accelerator and Brake: Dynamic Persuasion with Dead Ends

This paper studies dynamic persuasion in a strategic-experimentation relationship in which the principal has a single-peaked preference over the agent's stopping time. Excessive experimentation may end in a dead end. The principal privately observes project quality, which determines the agent's payoff conditional on success, while both parties learn about feasibility only through the agent's experimentation. We show that an optimal policy uses at most two one-shot disclosures: an accelerator before the principal's ideal stopping time and a brake afterward. A local Arrow--Pratt comparison of induced payoffs over stopping time determines whether the accelerator is concentrated or gradual. Under common discounting, the comparison yields a one-shot accelerator. Under heterogeneous discounting, the one-shot result remains robust unless the agent is sufficiently more impatient than the principal, in which case the ranking reverses over an interval and the accelerator can take a one-shot--gradual--one-shot form.

econ.TH

Modeling Human Behavior with Type Vectors Using AI

We introduce a general, easy-to-implement AI-based modeling technique for analyzing human behavior. A key feature of this approach, which contrasts with existing modeling techniques, is that it combines the flexibility and interpretability of natural language with a mathematical structure that can be fitted to data and easily analyzed. We assign a large language model a vector of trait intensities-a type vector-and then ask it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) could correspond to "You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5," after which it is asked to make choices. We can then vary the traits (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. We illustrate the method by applying it to model 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles. We find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. The type vectors needed to fit individuals across games cluster into fewer than a dozen groups, with substantial variation in fit across subjects. Moreover, the individual type vectors can predict behavior in held-out games with different rules and available actions. More broadly, this new modeling method is highly generalizable and interpretable: we can input any vector of traits and use them to model behavior across any setting

econ.TH