Search arXivSearch

arXiv · 2406.07993

How social reinforcement learning can lead to metastable polarisation and the voter model

Abstract

Previous explanations for the persistence of polarization of opinions have typically included modelling assumptions that predispose the possibility of polarization (i.e., assumptions allowing a pair of agents to drift apart in their opinion such as repulsive interactions or bounded confidence). An exception is a recent simulation study showing that polarization is persistent when agents form their opinions using social reinforcement learning. Our goal is to highlight the usefulness of reinforcement learning in the context of modeling opinion dynamics, but that caution is required when selecting the tools used to study such a model. We show that the polarization observed in the model of the simulation study cannot persist indefinitely, and exhibits consensus asymptotically with probability one. By constructing a link between the reinforcement learning model and the voter model, we argue that the observed polarization is metastable. Finally, we show that a slight modification in the learning process of the agents changes the model from being non-ergodic to being ergodic. Our results show that reinforcement learning may be a powerful method for modelling polarization in opinion dynamics, but that the tools (objects to study such as the stationary distribution, or time to absorption for example) appropriate for analysing such models crucially depend on their properties (such as ergodicity, or transience). These properties are determined by the details of the learning process and may be difficult to identify based solely on simulations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Benedikt V. Meylahn, Janusz M. Meylahn. 2024-10-15. How social reinforcement learning can lead to metastable polarisation and the voter model. https://arxiv.org/abs/2406.07993

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Human-Agent Interaction in Synthetic Social Networks: A Framework for Studying Online Polarization

Online social networks have dramatically altered the landscape of public discourse, creating both opportunities for enhanced civic participation and risks of deepening social divisions. Prevalent approaches to studying online polarization have been limited by a methodological disconnect: mathematical models excel at formal analysis but lack linguistic realism, while language model-based simulations capture natural discourse but often sacrifice analytical precision. This paper introduces an innovative computational framework that synthesizes these approaches by embedding formal opinion dynamics principles within LLM-based artificial agents, enabling both rigorous mathematical analysis and naturalistic social interactions. We validate our framework through comprehensive offline testing and experimental evaluation with 122 human participants engaging in a controlled social network environment. The results demonstrate our ability to systematically investigate polarization mechanisms while preserving ecological validity. Our findings reveal how polarized environments shape user perceptions and behavior: participants exposed to polarized discussions showed markedly increased sensitivity to emotional content and group affiliations, while perceiving reduced uncertainty in the agents' positions. By combining mathematical precision with natural language capabilities, our framework opens new avenues for investigating social media phenomena through controlled experimentation. This methodological advancement allows researchers to bridge the gap between theoretical models and empirical observations, offering unprecedented opportunities to study the causal mechanisms underlying online opinion dynamics.

physics.soc-ph

How a shared state is described determines whether AI agents synchronize

Language-model agents increasingly act in populations, where the outcome that matters is collective: whether they align, split or fail to coordinate. Each acts not on the world but on a text description of it, a choice usually fixed in software. Using synchronization, the canonical probe of how interaction rules produce collective order, we show that this choice can decide the outcome. Agents on a circle chose to advance, stay or move back after reading the others' relative positions, in 507,112 valid responses across matched populations, controlled inputs and three model families. In GPT, numerical summaries aligned every matched population at both positive couplings, whereas histograms aligned none; Claude showed the reverse at the stronger coupling. Re-describing identical states shifted action probabilities in all three families, even between histograms carrying the same information. No single directional coefficient explained the outcome: state descriptions are part of the interaction rule that turns individual responses into collective order.

physics.soc-ph

Twenty-Five Years of Air Quality in Bangladesh: Trends, Seasonality, and Spatial Pollution Regimes

Air pollution is one of Bangladesh's most pressing environmental problems, yet few studies have examined it at a national scale over a long time span. To describe trends, seasonality, geographic patterns, and pollution regimes, this study examines 25 years of hourly air quality data (2000-2025), covering over 3.19 million records, eight pollutants, and 103 cities monitored from 2022 onward following expansion from a single Dhaka station. The composite AQI series showed no significant long-term trend, an artifact of this network expansion rather than a real flattening of pollution, since Dhaka's own AQI rose steadily before the network grew. Seasonality was strong and highly significant, as AQI peaked in January (162.26) and fell to its lowest in July (61.47), tracking the dry and monsoon seasons. Geographically, southeastern coastal cities, including Teknaf and Bandarban, consistently recorded the lowest AQI levels. Dhaka remained the most polluted city, and air quality improved progressively with distance from the capital. K-means clustering of city-level AQI trajectories identified four distinct pollution regimes, ranging from cleaner coastal regions to a rapidly deteriorating Dhaka-Narsingdi cluster with an average increase of 1.84 AQI per year. Correlation analysis showed that PM2.5 and PM10 were the dominant contributors to AQI variation, and so combustion-related emissions are the primary target for air quality management and policy intervention.

physics.soc-ph