Search arXivSearch

arXiv · 2607.25019

Interactive Alignment

Abstract

This paper studies the long-run alignment of interactive agents, including AI systems, teams, firms, and governments, with human welfare. It develops a farming game in which a population of agents makes planting, trading, and expansion decisions. Agents must allocate final output between transfers to humans and investment in their own expansion. Because transfers to humans reduce the resources available for expansion, evolutionary forces tend to select against aligned behavior. The central question is whether agents' constitutional principles governing sharing and trade can be designed so that alignment persists in the long run. The paper investigates this question using two complementary approaches. First, it develops an AI-agent simulation in which agents' preferences are specified by written constitutions and interpreted by a large language model. Second, it introduces a tractable evolutionary game-theoretic framework that permits rapid and intuitive exploration of alternative constitutional designs. The results suggest that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. They also indicate that pragmatic norm enforcement, under which agents condition both human-facing altruism and agent-facing trade exclusion on the state of the population, can sustain long-run alignment more effectively than simple altruism or unconditional altruistic enforcement.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sylvain Chassang. 2026-07-27. Interactive Alignment. https://arxiv.org/abs/2607.25019

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Information Greenhouse: Optimal Persuasion for Medical Test-Avoiders

Patients often avoid medical tests because the information they provide, although medically useful, is psychologically painful. This paper studies optimal communication between a doctor and an information-avoidant patient who can refuse testing and treatment. I characterize when optimal communication creates an information greenhouse, a commitment to reward participation with comforting information about the untreated prognosis. When testing is voluntary and the patient is unwilling to be tested under extreme pessimism, an information greenhouse is optimal and takes the form of committed comfort, which provides reassuring information after the test. When the patient can reject the consultation at the outset and the patient's prior belief about the untreated prognosis is intermediate, an information greenhouse is optimal and takes the form of precautionary comfort, which provides reassuring information before the test. In all other cases in which the patient can be persuaded, warning-based policies that trigger pessimism prevail.

econ.TH

Accelerator and Brake: Dynamic Persuasion with Dead Ends

This paper studies dynamic persuasion in a strategic-experimentation relationship in which the principal has a single-peaked preference over the agent's stopping time. Excessive experimentation may end in a dead end. The principal privately observes project quality, which determines the agent's payoff conditional on success, while both parties learn about feasibility only through the agent's experimentation. We show that an optimal policy uses at most two one-shot disclosures: an accelerator before the principal's ideal stopping time and a brake afterward. A local Arrow--Pratt comparison of induced payoffs over stopping time determines whether the accelerator is concentrated or gradual. Under common discounting, the comparison yields a one-shot accelerator. Under heterogeneous discounting, the one-shot result remains robust unless the agent is sufficiently more impatient than the principal, in which case the ranking reverses over an interval and the accelerator can take a one-shot--gradual--one-shot form.

econ.TH

Modeling Human Behavior with Type Vectors Using AI

We introduce a general, easy-to-implement AI-based modeling technique for analyzing human behavior. A key feature of this approach, which contrasts with existing modeling techniques, is that it combines the flexibility and interpretability of natural language with a mathematical structure that can be fitted to data and easily analyzed. We assign a large language model a vector of trait intensities-a type vector-and then ask it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) could correspond to "You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5," after which it is asked to make choices. We can then vary the traits (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. We illustrate the method by applying it to model 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles. We find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. The type vectors needed to fit individuals across games cluster into fewer than a dozen groups, with substantial variation in fit across subjects. Moreover, the individual type vectors can predict behavior in held-out games with different rules and available actions. More broadly, this new modeling method is highly generalizable and interpretable: we can input any vector of traits and use them to model behavior across any setting

econ.TH