Search arXivSearch

arXiv · 2407.16034

Efficient Replay Memory Architectures in Multi-Agent Reinforcement Learning for Traffic Congestion Control

Abstract

Episodic control, inspired by the role of episodic memory in the human brain, has been shown to improve the sample inefficiency of model-free reinforcement learning by reusing high-return past experiences. However, the memory growth of episodic control is undesirable in large-scale multi-agent problems such as vehicle traffic management. This paper proposes a novel replay memory architecture called Dual-Memory Integrated Learning, to augment to multi-agent reinforcement learning methods for congestion control via adaptive light signal scheduling. Our dual-memory architecture mimics two core capabilities of human decision-making. First, it relies on diverse types of memory--semantic and episodic, short-term and long-term--in order to remember high-return states that occur often in the network and filter out states that don't. Second, it employs equivalence classes to group together similar state-action pairs and that can be controlled using the same action (i.e., light signal sequence). Theoretical analyses establish memory growth bounds, and simulation experiments on several intersection networks showcase improved congestion performance (e.g., vehicle throughput) from our method.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mukul Chodhary, Kevin Octavian, SooJean Han. 2024-07-22. Efficient Replay Memory Architectures in Multi-Agent Reinforcement Learning for Traffic Congestion Control. https://arxiv.org/abs/2407.16034

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cooperative Multi-Agent Assignment over Stochastic Graphs via Constrained Reinforcement Learning

Constrained multi-agent reinforcement learning offers the framework to design scalable and almost surely feasible solutions for teams of agents operating in dynamic environments to carry out conflicting tasks. We address the challenges of multi-agent coordination through an unconventional formulation in which the dual variables are not driven to convergence but are free to cycle, enabling agents to adapt their policies dynamically based on real-time constraint satisfaction levels. The coordination relies on a light single-bit communication protocol over a network with stochastic connectivity. Using this gossiped information, agents update local estimates of the dual variables. Furthermore, we modify the local dual dynamics by introducing a contraction factor, which lets us use finite communication buffers and keep the estimation error bounded. Under this model, we provide theoretical guarantees of almost sure feasibility and corroborate them with numerical experiments in which a team of robots successfully patrols multiple regions, communicating under a time-varying ad-hoc network.

eess.SY

Which Top Energy-Intensive Manufacturing Countries Can Compete in a Renewable Energy Future?

In a world increasingly powered by renewables and aiming for greenhouse gas-neutral industrial production, the future competitiveness of todays top manufacturing countries is questioned. This study applies detailed energy system modeling to quantify the Renewable Pull, an incentive for industry relocation exerted by countries with favorable renewable conditions. Results reveal that the Renewable Pull is not a cross-industrial phenomenon but strongly depends on the relationship between energy costs and transport costs. The intensity of the Renewable Pull varies, with China, India, and Japan facing a significantly stronger effect than Germany and the United States. Incorporating national capital cost assumptions proves critical, reducing Germanys Renewable Pull by a factor of six and positioning it as the second least affected top manufacturing country after Saudi Arabia. Using Germany as a case study, the analysis moreover illustrates that targeted import strategies, especially within the EU, can nearly eliminate the Renewable Pull, offering policymakers clear options for risk mitigation.

eess.SY

Certifying Frequency Stability for Systems with Line Dynamics and Heterogeneous Bus Dynamics

This work presents a framework for certifying small-signal frequency stability of a power system with line dynamics and heterogeneous bus dynamics. This framework can certify the stability of systems which include synchronous generators, synchronous condensers, and converter-interfaced resources with a wide range of controls. Moreover, it can do so without detailed or precise knowledge of the network topology. With this framework, we also provide a detailed analysis of how proportional-derivative (PD) droop can improve the stability margin of the frequency response. The stability certificates presented in this work, which extend prior results by incorporating line dynamics, provide insight into how the control parameters for different units in the system impact the overall frequency stability. While damper windings have long been understood to improve the frequency synchronization between machines, the dynamics of the damper windings are complex, making them difficult to analyze. To address this gap, this paper derives a novel reduced-order model of the damper windings in the form of a derivative droop term. Moreover, we show that derivative droop terms used in grid-forming (GFM) control can be understood as a form of damper winding emulation. Our analytical stability conditions highlight the importance of damper windings (or their emulation) in facilitating frequency synchronization and suppressing unstable interactions between GFM converters. These results are validated with electromagnetic-transient (EMT) simulation.

eess.SY