Search arXiv⌕ Search

arXiv · 2403.03224

Reinforcement Learning Jazz Improvisation: When Music Meets Game Theory

Abstract

Live performances of music are always charming, with the unpredictability of improvisation due to the dynamic between musicians and interactions with the audience. Jazz improvisation is a particularly noteworthy example for further investigation from a theoretical perspective. Here, we introduce a novel mathematical game theory model for jazz improvisation, providing a framework for studying music theory and improvisational methodologies. We use computational modeling, mainly reinforcement learning, to explore diverse stochastic improvisational strategies and their paired performance on improvisation. We find that the most effective strategy pair is a strategy that reacts to the most recent payoff (Stepwise Changes) with a reinforcement learning strategy limited to notes in the given chord (Chord-Following Reinforcement Learning). Conversely, a strategy that reacts to the partner's last note and attempts to harmonize with it (Harmony Prediction) strategy pair yields the lowest non-control payoff and highest standard deviation, indicating that picking notes based on immediate reactions to the partner player can yield inconsistent outcomes. On average, the Chord-Following Reinforcement Learning strategy demonstrates the highest mean payoff, while Harmony Prediction exhibits the lowest. Our work lays the foundation for promising applications beyond jazz: including the use of artificial intelligence (AI) models to extract data from audio clips to refine musical reward systems, and training machine learning (ML) models on existing jazz solos to further refine strategies within the game.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vedant Tapiavala, Joshua Piesner, Sourjyamoy Barman, Feng Fu. 2024-02-25. Reinforcement Learning Jazz Improvisation: When Music Meets Game Theory. https://arxiv.org/abs/2403.03224

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Allocation condensation under asymptotically complete mixing

We study a fixed network on which new mass is allocated in proportion to a power of each node's stock and then redistributed. Transport retains a fraction of both old and new mass at each node and divides the rest according to fixed background shares. We show that a small amount of retention can sustain concentration on networks where complete redistribution would leave every node's share of new mass vanishing as the network grows. A node receiving most new mass may retain little relative to the network total, yet much more than redistribution alone would supply. For a power above one, the allocation rule amplifies this relative stock advantage at the next step. We identify retention rates tending to zero with network size that allow widely spread and concentrated stationary allocations to coexist on the same network. One of them stays close to the allocation under complete redistribution. Each node can also dominate a separate concentrated allocation, receiving a share of new mass tending to one. The stocks supporting these allocations attract nearby trajectories, although every stationary stock approaches the same background, with even its largest share tending to zero. These allocations continue to coexist under specified changes in transport that preserve the background. Redistribution can therefore spread the stock widely without spreading new allocations in the same way.

physics.soc-ph↗

Glassy dynamics of metropolitan human mobility frozen near the gravitational equilibrium

This research establishes a connection between macroscopic urban commuting flows and thermal equilibrium. Using mobility data from around 30 million records across six Japanese cities over one year, we introduce the Gravity-based Home Swapping Model (GHSM). This model applies Metropolis dynamics to urban commuting by treating individuals as interacting particles. Within this framework, total commuting time dictates the system energy, and temperature controls how strongly mobility responds to cost savings. Consequently, urban commuting can be analyzed through an evaluable free energy and simulated similarly to physical systems of matter. The maximumentropy doubly constrained gravity model serves as the stationary state. We find that the residential dynamics exhibit glassy characteristics. Through calibration to observed data, we reveal that the system demonstrates signatures of kinetically constrained models, specifically ageing, hysteresis, and freezing near equilibrium. Simulations initialized from arbitrary states converge to empirical origin-destination matrices. This demonstrates that minimal physical mechanisms can reconstruct complex urban realities.

physics.soc-ph↗

Human-Agent Interaction in Synthetic Social Networks: A Framework for Studying Online Polarization

Online social networks have dramatically altered the landscape of public discourse, creating both opportunities for enhanced civic participation and risks of deepening social divisions. Prevalent approaches to studying online polarization have been limited by a methodological disconnect: mathematical models excel at formal analysis but lack linguistic realism, while language model-based simulations capture natural discourse but often sacrifice analytical precision. This paper introduces an innovative computational framework that synthesizes these approaches by embedding formal opinion dynamics principles within LLM-based artificial agents, enabling both rigorous mathematical analysis and naturalistic social interactions. We validate our framework through comprehensive offline testing and experimental evaluation with 122 human participants engaging in a controlled social network environment. The results demonstrate our ability to systematically investigate polarization mechanisms while preserving ecological validity. Our findings reveal how polarized environments shape user perceptions and behavior: participants exposed to polarized discussions showed markedly increased sensitivity to emotional content and group affiliations, while perceiving reduced uncertainty in the agents' positions. By combining mathematical precision with natural language capabilities, our framework opens new avenues for investigating social media phenomena through controlled experimentation. This methodological advancement allows researchers to bridge the gap between theoretical models and empirical observations, offering unprecedented opportunities to study the causal mechanisms underlying online opinion dynamics.

physics.soc-ph↗