Search arXiv⌕ Search

arXiv · 2009.04350

Reinforcement Learning in Non-Stationary Discrete-Time Linear-Quadratic Mean-Field Games

Abstract

In this paper, we study large population multi-agent reinforcement learning (RL) in the context of discrete-time linear-quadratic mean-field games (LQ-MFGs). Our setting differs from most existing work on RL for MFGs, in that we consider a non-stationary MFG over an infinite horizon. We propose an actor-critic algorithm to iteratively compute the mean-field equilibrium (MFE) of the LQ-MFG. There are two primary challenges: i) the non-stationarity of the MFG induces a linear-quadratic tracking problem, which requires solving a backwards-in-time (non-causal) equation that cannot be solved by standard (causal) RL algorithms; ii) Many RL algorithms assume that the states are sampled from the stationary distribution of a Markov chain (MC), that is, the chain is already mixed, an assumption that is not satisfied for real data sources. We first identify that the mean-field trajectory follows linear dynamics, allowing the problem to be reformulated as a linear quadratic Gaussian problem. Under this reformulation, we propose an actor-critic algorithm that allows samples to be drawn from an unmixed MC. Finite-sample convergence guarantees for the algorithm are then provided. To characterize the performance of our algorithm in multi-agent RL, we have developed an error bound with respect to the Nash equilibrium of the finite-population game.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Muhammad Aneeq uz Zaman, Kaiqing Zhang, Erik Miehling, Tamer Başar. 2020-10-01. Reinforcement Learning in Non-Stationary Discrete-Time Linear-Quadratic Mean-Field Games. https://arxiv.org/abs/2009.04350

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Data to Sliding Mode Control of Uncertain Large-Scale Networks with Unknown Dynamics

In this paper, we develop a compositional data-driven approach for the global stabilization of large-scale nonlinear networks with unknown dynamics and external perturbations. We first collect data along a single trajectory of each unknown nominal subsystem during a finite-time experiment. The data collected from each nominal subsystem are then used to design a feedback law that renders each nominal closed-loop subsystem input-to-state stable (ISS), certified by its corresponding ISS Lyapunov function. We derive conditions as data-dependent semidefinite programs that simultaneously yield local ISS controllers and the corresponding ISS Lyapunov functions. To cancel the effect of external perturbations on subsystem dynamics and, consequently, on the whole network dynamics, we then design a local integral sliding mode (ISM) controller for each subsystem using the collected data. Under a small-gain compositional condition, we employ data-driven ISS Lyapunov functions designed for the subsystems to construct a control Lyapunov function for the network, guaranteeing that the nominal closed-loop network is globally asymptotically stable (GAS) at the origin. We then extend this compositional result to perturbed networks, proving that the synthesized ISM controllers render the origin of the closed-loop network GAS even in the presence of perturbations. We demonstrate the efficacy of the proposed data-driven approach on large-scale interconnected networks with five distinct interconnection topologies.

eess.SY↗

Simultaneous improvement of control and estimation for battery management systems

Standard battery management systems treat the control and state estimation problems as decoupled objectives, relying on certainty equivalence controllers that are blind to the varying observability induced by nonlinear open-circuit voltage models. In this paper, we show that for a broad class of objectives, including the peak shaving and valley filling scenarios common in grid-connected energy storage, the expected cost of a stochastic battery system can be exactly parametrized by the conditional mean and covariance of the state of charge. This reformulation reveals a direct coupling between the control input and estimation quality, a coupling that certainty equivalence controllers ignore, and motivates a dual-control approach in which the controller actively reduces estimation uncertainty by driving the state to high observability regions without compromising the control objective. We derive a deterministic surrogate to this stochastic cost and pose the dual-control problem as a computationally tractable model predictive control problem. We validate our approach on a nine-battery system tracking a time-varying reference trajectory. We report simultaneous improvements in tracking cost (a 28\% reduction) and state estimation error (up to 18\% reduction). The estimation improvement is reported across different state estimators: extended Kalman filter, unscented Kalman filter, and a moving horizon estimator, confirming that the estimation improvement of our approach is not restricted to a specific state observer.

eess.SY↗

Time-To-Reach Separation and Safety Filtering for Safe, Fair, and Efficient Multi-Agent Coordination

Advanced Air Mobility operations are expected to significantly increase aerial traffic in urban airspace, requiring autonomous traffic management systems to ensure collision-free operations in highly congested environments. In this paper, we propose a multi-agent coordination framework that uses minimum time-to-reach (TTR) as a unifying metric for priority assignment, temporal separation, and safety filtering. We focus on the problem of coordinating multiple aerial vehicles merging into an air corridor while maintaining safe separation between vehicles. Vehicles are assigned arrival-consistent priority based on TTR, and target TTR values are used to enforce temporal spacing, which induces spatial separation. A priority-consistent safety filtering layer based on Hamilton-Jacobi reachability value functions promotes collision avoidance while minimally modifying the reference guidance. Simulation results in a highly congested corridor merging scenario show that the proposed method improves safety, fairness, and efficiency compared to time-optimal guidance and priority-agnostic safety filtering.

eess.SY↗