Search arXiv⌕ Search

arXiv · 2610.01667

Conditioning LLMs on Social Value Orientation improves behavioural alignment in a sequential social dilemma

Abstract

Large language Models (LLMs) are increasingly used to simulate human decision-making, yet their outputs often under-represent human behavioural heterogeneity. We investigate whether conditioning LLMs on Social Value Orientation (SVO, a measure of how individuals value their own outcomes relative to others') can better reproduce human behaviour in a sequential social dilemma. Using experimental data from two variants of the Centipede Game (CG) as reference, we compare the default behaviour of eight LLMs with behaviour generated after conditioning them on SVO profiles drawn from the human sample. We find that the default strategies vary substantially across models, but are generally distant from the reference distributions. However, conditioning them on SVO profiles systematically steers their strategic behaviour, improving alignment with the human reference by up to 70% with respect to the default. Across models, higher induced SVO values decrease the probability of stopping the game, reproducing the relationship between prosociality and cooperation observed in human behaviour. Together with the sensitivity of LLMs' elicited SVO to prompt and order effects, these results suggest that the usefulness of SVO for behavioural simulation does not depend on LLMs possessing stable social preferences, but rather on their ability to map social preferences onto corresponding strategic choices.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marco Saponara, Axel Abels, Ann Nowé, Tom Lenaerts. 2026-10-01. Conditioning LLMs on Social Value Orientation improves behavioural alignment in a sequential social dilemma. https://arxiv.org/abs/2610.01667

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Perspectives on Unsolvability in the Roommates Problem

Instances of the well-studied Stable Roommates problem need not admit stable matchings. A long-standing open question posed by Gusfield and Irving (1989) asks about the behaviour of the function Pn, which measures the likelihood that a random instance with n agents is solvable (i.e., admits at least one stable matching). While very recently resolved in the limit for the case where n is even and preferences are sampled uniformly at random, this paper provides a comprehensive analysis of the landscape surrounding this question, combining structural, probabilistic, and experimental perspectives. We estimate Pn for instances with preferences sampled from diverse statistical distributions, for even and odd numbers of agents, examining problem sizes up to 5,001 agents, and considering important substructures. Our results reveal that while Pn tends to be low for most distributions, the number and lengths of "unstable" structures remain limited, suggesting that random instances are "close" to being solvable. Additionally, we present the first empirical study of the number of stable matchings and partitions that random instances admit. Our findings show that the solution sets are typically small, which suggests that many NP-hard problems related to computing optimal stable matchings and partitions become tractable in practice.

cs.GT↗

Fair and Efficient Investment in Public Transportation

We study a stylized model of infrastructure investment in public transportation. In our model, each agent travels between a pair of terminals in a network captured by a weighted graph, where edge weights represent distances. The central planner can reduce the travel time along a fixed number of edges, with the goal of maximizing the utilitarian or egalitarian welfare. When there is only one agent, we provide a polynomial-time algorithm that combines Dijkstra's algorithm with a dynamic program. We then demonstrate how to use this algorithm as a subroutine to solve the problem for two agents. Generalizing this idea, we present an XP algorithm parameterized by the number of agents $n$; however, our problem turns out to be W[1]-hard with respect to $n$. Nevertheless, we establish a fixed-parameter tractability result for the special case where all agents travel to a common hub. If the number of agents is variable, we obtain NP-completeness and inapproximability results. We discuss implications of our results for a related model of railway network design.

cs.GT↗

Online Generalized-Mean Welfare Maximization: Achieving Near-Optimal Regret from Samples

We study online fair allocation of $T$ sequentially arriving items among $n$ agents with heterogeneous preferences, with the objective of maximizing generalized-mean welfare, defined as the $p$-mean of agents' time-averaged utilities, with $p\in (-\infty, 1)$. We first consider the i.i.d. arrival model and show that the pure greedy algorithm -- which myopically chooses the welfare-maximizing integral allocation -- achieves $\widetilde{O}(1/T)$ average regret. Importantly, in contrast to prior work, our algorithm does not require distributional knowledge and achieves the optimal regret rate using only the online samples. We then go beyond i.i.d. arrivals and investigate a nonstationary model with time-varying independent distributions. In the absence of additional data about the distributions, it is known that every online algorithm must suffer $Ω(1)$ average regret. We show that only a single historical sample from each distribution is sufficient to recover the optimal $\widetilde{O}(1/T)$ average regret rate, even in the face of arbitrary non-stationarity. Our algorithms are based on the re-solving paradigm: they assume that the remaining items will be the ones seen historically in those periods and solve the resulting welfare-maximization problem to determine the decision in every period. Finally, we also account for distribution shifts that may distort the fidelity of historical samples and show that the performance of our re-solving algorithms is robust to such shifts.

cs.GT↗