Search arXivSearch

arXiv subjects

Daniel Lacker

Publications and source records attributed to Daniel Lacker.

At least 19 recordsLinked to original sources

The Mean-Field Limit of Online Stochastic Vector Balancing

We study an online vector balancing problem, in which $n$ independent Gaussian random vectors $\boldsymbol{\zeta}(1),\dots,\boldsymbol{\zeta}(n) \sim \mathcal{N}(0, I_n)$, each of dimension $n$, arrive one at a time. The goal is to choose signs $\varepsilon(1),\dots,\varepsilon(n) \in \{\pm 1\}$ with $\varepsilon(k)$ depending only on $\boldsymbol{\zeta}(1),\dots,\boldsymbol{\zeta}(k)$, so as to minimize the expected $\ell^{\infty}$ norm of the signed sum $\frac{1}{\sqrt{n}}\sum_{k = 1}^n \varepsilon(k) \boldsymbol{\zeta}(k)$. Prior work showed that the optimal value $V^n$ is $O(1)$, at least for Rademacher $\boldsymbol{\zeta}(k)$'s, by constructing specific algorithms. Our main contribution is to determine the exact limit $V^{\infty} = \lim_{n\to\infty} V^n$ as the value of a nonstandard stochastic control problem of mean-field type: find the narrowest terminal interval into which a Brownian motion can be adaptively steered under a uniform-in-time $L^2$ constraint on the drift. The proof of the lower bound $V^{\infty} \leq \liminf_{n \to \infty} V^n$ uses probabilistic compactness arguments, and is very flexible. In fact, we show that the lower bound is universal, in that it holds as long as the entries of the $\boldsymbol{\zeta}(k)$ vectors are i.i.d. with mean zero, variance 1, and finite fourth moment. The proof of the upper bound $\limsup_{n \to \infty} V^n \leq V^{\infty}$ is more delicate, relying on dynamic programming principles and a priori bounds obtained from a coupling procedure involving the F\"ollmer drift, which makes explicit use of the Gaussian structure. In addition to our main convergence result, we provide some analysis and asymptotics for the limiting mean-field control problem.

math.PR

Sharp propagation of chaos for mean field Langevin dynamics, control, and games

We establish the sharp rate of propagation of chaos for McKean-Vlasov equations with coefficients that are non-linear in the measure argument, i.e., not necessarily given by pairwise interactions. Results are given both on bounded time horizon and uniform in time. As applications, we deduce the sharp rate of propagation of chaos for the convergence problem in mean field games and control, and for mean field Langevin dynamics, the latter being uniform in time in the strongly displacement convex regime. Our arguments combine the BBGKY hierarchy with techniques from the literature on weak propagation of chaos.

math.PR

Marginal dynamics of probabilistic cellular automata on trees

We study locally interacting processes in discrete time, often called probabilistic cellular automata, indexed by locally finite graphs. For infinite regular trees and certain generalized Galton-Watson trees, we show that the marginal evolution at a single vertex and its neighborhood can be characterized by an autonomous stochastic recursion referred to as the local-field equation. This evolution can be viewed as a nonlinear or measure-dependent chain, but the measure dependence arises from the symmetries of the underlying tree rather than from any mean field interactions. We discuss applications to simulation of marginal dynamics and approximations of empirical measures of interacting chains on several generic classes of large-scale finite graphs that are locally tree-like. In addition to the symmetries of the tree, a key role is played by a second-order Markov random field property, which we establish for general graphs along with some other novel Gibbs measure properties.

math.PR

A hierarchical entropy method for the delocalization of bias in high-dimensional Langevin Monte Carlo

The unadjusted Langevin algorithm is widely used for sampling from complex high-dimensional distributions. It is well known to be biased, with the bias typically scaling linearly with the dimension when measured in squared Wasserstein distance. However, the recent paper of Chen et al. (2024) identifies an intriguing new delocalization effect: For a class of distributions with sparse interactions, the bias between low-dimensional marginals scales only with the lower dimension, not the full dimension. In this work, we strengthen the results of Chen et al. (2024) in the sparse interaction regime by removing a logarithmic factor, measuring distance in relative entropy (a.k.a. KL-divergence), and relaxing the strong log-concavity assumption. In addition, we expand the scope of the delocalization phenomenon by showing that it holds for a class of distributions with weak interactions. Our proofs are based on a hierarchical analysis of the marginal relative entropies, inspired by the authors' recent work on propagation of chaos.

stat.ML

Geodesic convexity and strengthened functional inequalities in submanifolds of Wasserstein space

We study the geodesic convexity of various energy and entropy functionals restricted to (non-geodesically convex) submanifolds of Wasserstein spaces with their induced geometry. We prove a variety of convexity results by means of a simple general principle, which holds in the metric space setting, and which crucially requires no knowledge of the structure of geodesics in the submanifold: If the EVI gradient flow of a functional exists and leaves the submanifold invariant, then the restriction of the functional to the submanifold is geodesically convex. This leads to short new proofs of several known results, such as one of Carlen and Gangbo on strong convexity of entropy on sphere-like submanifolds, and several new results, such as the $\lambda$-convexity of entropy on the space of couplings of $\lambda$-log-concave marginals. Along the way, we develop sufficient conditions for existence of geodesics in Wasserstein submanifolds. Submanifold convexity results lead systematically to improvements of Talagrand and HWI inequalities which we speculate to be closely related to concentration of measure estimates for conditioned empirical measures, and we prove one rigorous result in this direction in the Carlen-Gangbo setting.

math.AP

Mimicking and Conditional Control with Hard Killing

We first prove a mimicking theorem (also known as a Markovian projection theorem) for the marginal distributions of an Ito process conditioned to not have exited a given domain. We then apply this new result to the proof of a conjecture of P.L. Lions for the optimal control of conditioned processes.

math.PR

Quantitative propagation of chaos for non-exchangeable diffusions via first-passage percolation

This paper develops a non-asymptotic approach to mean field approximations for systems of $n$ diffusive particles interacting pairwise. The interaction strengths are not identical, making the particle system non-exchangeable. The marginal law of any subset of particles is compared to a suitably chosen product measure, and we find sharp relative entropy estimates between the two. Building upon prior work of the first author in the exchangeable setting, we use a generalized form of the BBGKY hierarchy to derive a hierarchy of differential inequalities for the relative entropies. Our analysis of this complicated hierarchy exploits an unexpected but crucial connection with first-passage percolation, which lets us bound the marginal entropies in terms of expectations of functionals of this percolation process.

math.PR

Convergence of coordinate ascent variational inference for log-concave measures via optimal transport

Mean field variational inference (VI) is the problem of finding the closest product (factorized) measure, in the sense of relative entropy, to a given high-dimensional probability measure $\rho$. The well known Coordinate Ascent Variational Inference (CAVI) algorithm aims to approximate this product measure by iteratively optimizing over one coordinate (factor) at a time, which can be done explicitly. Despite its popularity, the convergence of CAVI remains poorly understood. In this paper, we prove the convergence of CAVI for log-concave densities $\rho$. If additionally $\log \rho$ has Lipschitz gradient, we find a linear rate of convergence, and if also $\rho$ is strongly log-concave, we find an exponential rate. Our analysis starts from the observation that mean field VI, while notoriously non-convex in the usual sense, is in fact displacement convex in the sense of optimal transport when $\rho$ is log-concave. This allows us to adapt techniques from the optimization literature on coordinate descent algorithms in Euclidean space.

stat.ML

Independent projections of diffusions: Gradient flows for variational inference and optimal mean field approximations

What is the optimal way to approximate a high-dimensional diffusion process by one in which the coordinates are independent? This paper presents a construction, called the \emph{independent projection}, which is optimal for two natural criteria. First, when the original diffusion is reversible with invariant measure $\rho_*$, the independent projection serves as the Wasserstein gradient flow for the relative entropy $H(\cdot\,|\,\rho_*)$ constrained to the space of product measures. This is related to recent Langevin-based sampling schemes proposed in the statistical literature on mean field variational inference. In addition, we provide both qualitative and quantitative results on the long-time convergence of the independent projection, with quantitative results in the log-concave case derived via a new variant of the logarithmic Sobolev inequality. Second, among all processes with independent coordinates, the independent projection is shown to exhibit the slowest growth rate of path-space entropy relative to the original diffusion. This sheds new light on the classical McKean-Vlasov equation and recent variants proposed for non-exchangeable systems, which can be viewed as special cases of the independent projection.

math.PR

Projected Langevin dynamics and a gradient flow for entropic optimal transport

The classical (overdamped) Langevin dynamics provide a natural algorithm for sampling from its invariant measure, which uniquely minimizes an energy functional over the space of probability measures, and which concentrates around the minimizer(s) of the associated potential when the noise parameter is small. We introduce analogous diffusion dynamics that sample from an entropy-regularized optimal transport, which uniquely minimizes the same energy functional but constrained to the set $\Pi(\mu,\nu)$ of couplings of two given marginal probability measures $\mu$ and $\nu$ on $\mathbb{R}^d$, and which concentrates around the optimal transport coupling(s) for small regularization parameter. More specifically, our process satisfies two key properties: First, the law of the solution at each time stays in $\Pi(\mu,\nu)$ if it is initialized there. Second, the long-time limit is the unique solution of an entropic optimal transport problem. In addition, we show by means of a new log-Sobolev-type inequality that the convergence holds exponentially fast, for sufficiently large regularization parameter and for a class of marginals which strictly includes all strongly log-concave measures. By studying the induced Wasserstein geometry of the submanifold $\Pi(\mu,\nu)$, we argue that the SDE can be viewed as a Wasserstein gradient flow on this space of couplings, at least when $d=1$, and we identify a conjectural gradient flow for $d \ge 2$. The main technical difficulties stems from the appearance of conditional expectation terms which serve to constrain the dynamics to $\Pi(\mu,\nu)$.

math.PR

Approximately optimal distributed stochastic controls beyond the mean field setting

We study high-dimensional stochastic optimal control problems in which many agents cooperate to minimize a convex cost functional. We consider both the full-information problem, in which each agent observes the states of all other agents, and the distributed problem, in which each agent observes only its own state. Our main results are sharp non-asymptotic bounds on the gap between these two problems, measured both in terms of their value functions and optimal states. Along the way, we develop theory for distributed optimal stochastic control in parallel with the classical setting, by characterizing optimizers in terms of an associated stochastic maximum principle and a Hamilton-Jacobi-type equation. By specializing these results to the setting of mean field control, in which costs are (symmetric) functions of the empirical distribution of states, we derive the optimal rate for the convergence problem in the displacement convex regime.

math.PR

Mean field approximations via log-concavity

We propose a new approach to deriving quantitative mean field approximations for any probability measure $P$ on $\mathbb{R}^n$ with density proportional to $e^{f(x)}$, for $f$ strongly concave. We bound the mean field approximation for the log partition function $\log \int e^{f(x)}dx$ in terms of $\sum_{i \neq j}\mathbb{E}_{Q^*}|\partial_{ij}f|^2$, for a semi-explicit probability measure $Q^*$ characterized as the unique mean field optimizer, or equivalently as the minimizer of the relative entropy $H(\cdot\,|\,P)$ over product measures. This notably does not involve metric-entropy or gradient-complexity concepts which are common in prior work on nonlinear large deviations. Three implications are discussed, in the contexts of continuous Gibbs measures on large graphs, high-dimensional Bayesian linear regression, and the construction of decentralized near-optimizers in high-dimensional stochastic control problems. Our arguments are based primarily on functional inequalities and the notion of displacement convexity from optimal transport.

math.PR

Sharp uniform-in-time propagation of chaos

We prove the optimal rate of quantitative propagation of chaos, uniformly in time, for interacting diffusions. Our main examples are interactions governed by convex potentials and models on the torus with small interactions. We show that the distance between the $k$-particle marginal of the $n$-particle system and its limiting product measure is $O((k/n)^2)$, uniformly in time, with distance measured either by relative entropy, squared quadratic Wasserstein metric, or squared total variation. Our proof is based on an analysis of relative entropy through the BBGKY hierarchy, adapting prior work of the first author to the time-uniform case by means of log-Sobolev inequalities.

math.PR

A label-state formulation of stochastic graphon games and approximate equilibria on large networks

This paper studies stochastic games on large graphs and their graphon limits. We propose a new formulation of graphon games based on a single typical player's label-state distribution. In contrast, other recently proposed models of graphon games work directly with a continuum of players, which involves serious measure-theoretic technicalities. In fact, by viewing the label as a component of the state process, we show in our formulation that graphon games are a special case of mean field games, albeit with certain inevitable degeneracies and discontinuities that make most existing results on mean field games inapplicable. Nonetheless, we prove existence of Markovian graphon equilibria under fairly general assumptions, as well as uniqueness under a monotonicity condition. Most imporantly, we show how our notion of graphon equilibrium can be used to construct approximate equilibria for large finite games set on any (weighted, directed) graph which converges in cut norm. The lack of players' exchangeability necessitates a careful definition of approximate equilibrium, allowing heterogeneity among the players' approximation errors, and we show how various regularity properties of the model inputs and underlying graphon lead naturally to different strengths of approximation.

math.OC

Stationary solutions and local equations for interacting diffusions on regular trees

We study the invariant measures of infinite systems of stochastic differential equations (SDEs) indexed by the vertices of a regular tree. These invariant measures correspond to Gibbs measures associated with certain continuous specifications, and we focus specifically on those measures which are homogeneous Markov random fields. We characterize the joint law at any two adjacent vertices in terms of a new two-dimensional SDE system, called the "local equation", which exhibits an unusual dependence on a conditional law. Exploiting an alternative characterization in terms of an eigenfunction-type fixed point problem, we derive existence and uniqueness results for invariant measures of the local equation and infinite SDE system. This machinery is put to use in two examples. First, we give a detailed analysis of the surprisingly subtle case of linear coefficients, which yields a new way to derive the famous Kesten-McKay law for the spectral measure of the regular tree. Second, we construct solutions of tree-indexed SDE systems with nearest-neighbor repulsion effects, similar to Dyson's Brownian motion.

math.PR

Closed-loop convergence for mean field games with common noise

This paper studies the convergence problem for mean field games with common noise. We define a suitable notion of weak mean field equilibria, which we prove captures all subsequential limit points, as $n\to\infty$, of closed-loop approximate equilibria from the corresponding $n$-player games. This extends to the common noise setting a recent result of the first author, while also simplifying a key step in the proof and allowing unbounded coefficients and non-i.i.d. initial conditions. Conversely, we show that every weak mean field equilibrium arises as the limit of some sequence of approximate equilibria for the $n$-player games, as long as the latter are formulated over a broader class of closed-loop strategies which may depend on an additional common signal.

math.PR

Quantitative approximate independence for continuous mean field Gibbs measures

Many Gibbs measures with mean field interactions are known to be chaotic, in the sense that any collection of $k$ particles in the $n$-particle system are asymptotically independent, as $n\to\infty$ with $k$ fixed or perhaps $k=o(n)$. This paper quantifies this notion for a class of continuous Gibbs measures on Euclidean space with pairwise interactions, with main examples being systems governed by convex interactions and uniformly convex confinement potentials. The distance between the marginal law of $k$ particles and its limiting product measure is shown to be $O((k/n)^{c \wedge 2})$, with $c$ proportional to the squared temperature. In the high temperature case, this improves upon prior results based on subadditivity of entropy, which yield $O(k/n)$ at best. The bound $O((k/n)^2)$ cannot be improved, as a Gaussian example demonstrates. The results are non-asymptotic, and distance is quantified via relative Fisher information, relative entropy, or the squared quadratic Wasserstein metric. The method relies on an a priori functional inequality for the limiting measure, used to derive an estimate for the $k$-particle distance in terms of the $(k+1)$-particle distance.

math.PR

Hierarchies, entropy, and quantitative propagation of chaos for mean field diffusions

This paper develops a non-asymptotic, local approach to quantitative propagation of chaos for a wide class of mean field diffusive dynamics. For a system of $n$ interacting particles, the relative entropy between the marginal law of $k$ particles and its limiting product measure is shown to be $O((k/n)^2)$ at each time, as long as the same is true at time zero. A simple Gaussian example shows that this rate is optimal. The main assumption is that the limiting measure obeys a certain functional inequality, which is shown to encompass many potentially irregular but not too singular finite-range interactions, as well as some infinite-range interactions. This unifies the previously disparate cases of Lipschitz versus bounded measurable interactions, improving the best prior bounds of $O(k/n)$ which were deduced from global estimates involving all $n$ particles. We also cover a class of models for which qualitative propagation of chaos and even well-posedness of the McKean-Vlasov equation were previously unknown. At the center of a new approach is a differential inequality, derived from a form of the BBGKY hierarchy, which bounds the $k$-particle entropy in terms of the $(k+1)$-particle entropy.

math.PR