Search arXiv⌕ Search

arXiv · 2507.01420

Reinforcement Learning for Discrete-time LQG Mean Field Social Control Problems with Unknown Dynamics

Abstract

This paper studies the discrete-time linear-quadratic-Gaussian mean field (MF) social control problem in an infinite horizon, where the dynamics of all agents are unknown. The objective is to design a reinforcement learning (RL) algorithm to approximate the decentralized asymptotic optimal social control in terms of two algebraic Riccati equations (AREs). In this problem, a coupling term is introduced into the system dynamics to capture the interactions among agents. This causes the equivalence between model-based and model-free methods to be invalid, which makes it difficult to directly apply traditional model-free algorithms. Firstly, under the assumptions of system stabilizability and detectability, a model-based policy iteration algorithm is proposed to approximate the stabilizing solution of the AREs. The algorithm is proven to be convergent in both cases of semi-positive definite and indefinite weight matrices. Subsequently, by adopting the method of system transformation, a model-free RL algorithm is designed to solve for asymptotic optimal social control. During the iteration process, the updates are performed using data collected from any two agents and MF state. Finally, a numerical case is provided to verify the effectiveness of the proposed algorithm.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hanfang Zhang, Bing-Chang Wang, Shuo Chen. 2025-12-04. Reinforcement Learning for Discrete-time LQG Mean Field Social Control Problems with Unknown Dynamics. https://arxiv.org/abs/2507.01420

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Symmetry-dependence in Rounding of a Convex Body

The symmetry measure of a convex body $S\subset\mathbb{R}^n$ is given by: $\mathrm{sym}(S):=\max\{α\ge0:\text{ there exists }x\in S\text{ such that }-α(S-x)\subseteq S-x\}$, where such an $x$ is called a Minkowski center. We prove that every convex body $S$ admits a $\sqrt{\frac{n}{\mathrm{sym}(S)}}$-rounding of $S$, namely, there exists an origin-centered ellipsoid $E$ and a center $c$ such that $E\subseteq S-c\subseteq\sqrt{\frac{n}{\mathrm{sym}(S)}}\,E$. This result was conjectured in 2005 by Belloni and Freund. As special cases, this recovers an $n$-rounding of $S$ (since $\mathrm{sym}(S)\ge\frac{1}{n}$), and a $\sqrt{n}$-rounding when $\mathrm{sym}(S)=1$. In the case when $S$ is a polytope given as the convex hull of points, the desired rounding is produced by a regularized minimum-volume covering ellipsoid problem where the regularization is with respect to the Minkowski center. Similarly, when $S$ is a polytope given as the intersection of halfspaces, such a rounding is produced by a regularized maximum-volume inscribed ellipsoid problem. In both of these cases, the rounding can be computed by first solving a linear optimization problem (to compute $\mathrm{sym}(S)$ and a Minkowski center), and then solving a convex optimization problem with a logarithmic determinant objective, second-order cone constraints, and one semidefinite cone constraint. We also show that the factor $\sqrt{\frac{n}{\mathrm{sym}(S)}}$ is nearly tight in its dependence on dimension and symmetry. When $\frac{n+1}{1+\mathrm{sym}(S)}$ is an integer, we show by explicit construction that the factor $\sqrt{\frac{n}{\mathrm{sym}(S)}}$ is tight. In the more general case, for every dimension $n$ and every admissible symmetry value, we construct a polytope $S$ for which every rounding factor is at least $\sqrt{\frac{2}{3}}\sqrt{\frac{n}{\mathrm{sym}(S)}}$.

math.OC↗

Maximal Monotone Differential Inclusions with Volterra and One-Sided Lipschitz Perturbations under Nonlocal Initial Conditions and Applications

We investigate a class of differential inclusions governed by non-autonomous and autonomous maximal monotone operators, involving set-valued perturbations with a Volterra integral term and subject to a nonlocal condition. Under suitable assumptions on the governing operator, the set-valued perturbation, and the Volterra kernel, we establish existence results for solutions. In particular, the set-valued perturbation is assumed to satisfy a one-sided Lipschitz condition, while the Volterra kernel is required to be Lipschitz continuous. The existence of a trajectory is established by means of an iterative construction and an application of Zorn's lemma. Finally, several examples are presented to illustrate the applicability of the abstract results.

math.OC↗

On Fast-Slow Mean-Field Forward-Backward Stochastic Systems

We establish an averaging principle for a class of multiscale mean-field forward-backward stochastic differential equations and identify several novel phenomena that are absent from classical fast-slow systems. In contrast with classical fast-slow systems, the effective dynamics cannot in general be obtained by simply freezing deterministic slow parameters and averaging against the invariant measure of the resulting fast equation. The appropriate averaging object is instead provided by a frozen fast dynamics in a random environment and its associated conditional invariant measures, which retain the coupling between the slow state and its distribution. The forward-backward structure creates a further obstruction: local averaging estimates need not remain stable when propagated over an arbitrary time horizon. We identify a uniform restart stability condition for the averaged system under which this obstruction can be overcome. Using a joint lifted semigroup for the state-law dynamics, together with a two-scale discretization and a Gordin-type decomposition, we prove strong averaging for both the forward and backward components with optimal convergence rate $O(\varepsilon^{1/2})$. As an application, we apply the general theory to a class of mean-field stochastic control problems and develop an efficient algorithm for solving such mean-field control problems.

math.OC↗