Search arXivSearch

arXiv · 1909.09417

Regularized Diffusion Adaptation via Conjugate Smoothing

Abstract

The purpose of this work is to develop and study a distributed strategy for Pareto optimization of an aggregate cost consisting of regularized risks. Each risk is modeled as the expectation of some loss function with unknown probability distribution while the regularizers are assumed deterministic, but are not required to be differentiable or even continuous. The individual, regularized, cost functions are distributed across a strongly-connected network of agents and the Pareto optimal solution is sought by appealing to a multi-agent diffusion strategy. To this end, the regularizers are smoothed by means of infimal convolution and it is shown that the Pareto solution of the approximate, smooth problem can be made arbitrarily close to the solution of the original, non-smooth problem. Performance bounds are established under conditions that are weaker than assumed before in the literature, and hence applicable to a broader class of adaptation and learning problems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stefan Vlaski, Lieven Vandenberghe, Ali H. Sayed. 2019-09-20. Regularized Diffusion Adaptation via Conjugate Smoothing. https://arxiv.org/abs/1909.09417

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stochastic Optimization Algorithms for Problems with Controllable Biased Oracles

Motivated by emerging applications in machine learning, we consider an optimization problem in a general setting in which the gradient of the objective function is available via a biased stochastic oracle. We assume a bias-control parameter can reduce the bias magnitude; however, a lower bias requires more computation/samples. For instance, in two applications on stochastic composition optimization and policy optimization for infinite-horizon Markov decision processes, we show that the bias follows a power law and exponential decay, respectively, as functions of their corresponding bias control parameters. For problems with such gradient oracles, the paper proposes stochastic algorithms that adjust the bias-control parameter throughout the iterations. We analyze the nonasymptotic performance of the proposed algorithms in the nonconvex regime and establish their sample or bias-control computation complexities to obtain a stationary point in expectation or with high probability. Finally, we numerically evaluate the performance of the proposed algorithms over three applications.

math.OC

Inverse Problems Over Probability Measure Space

Define a forward problem as $ρ_y = G_\#ρ_x$, where the probability distribution $ρ_x$ is mapped to another distribution $ρ_y$ using the forward operator $G$. In this work, we investigate the corresponding inverse problem: Given $ρ_y$, how to find $ρ_x$? Depending on whether $ G$ is overdetermined or underdetermined, the solution can have drastically different behavior. In the overdetermined case, we formulate a variational problem $\min_{ρ_x} D( G_\#ρ_x, ρ_y)$, and find that different choices of the metric $ D$ significantly affect the quality of the reconstruction. When $ D$ is set to be the Wasserstein distance, the reconstruction is the marginal distribution, while setting $ D$ to be a $ϕ$-divergence reconstructs the conditional distribution. In the underdetermined case, we formulate the constrained optimization $\min_{\{ G_\#ρ_x=ρ_y\}} E[ρ_x]$. The choice of $ E$ also significantly impacts the construction: setting $ E$ to be the entropy gives us the piecewise constant reconstruction, while setting $ E$ to be the second moment, we recover the classical least-norm solution. We also examine the formulation with regularization: $\min_{ρ_x} D( G_\#ρ_x, ρ_y) + α\mathsf R[ρ_x]$, and find that the entropy-entropy pair leads to a regularized solution that is defined in a piecewise manner, whereas the $W_2$-$W_2$ pair leads to a least-norm solution where $W_2$ is the 2-Wasserstein metric.

math.OC

When Does Selfishness Align with Team Goals? A Structural Analysis of Equilibrium and Optimality

This paper investigates the relationship between the team-optimal solution and the Nash equilibrium (NE) to assess the impact of self-interested decisions on team performance. In classical team decision problems, team members typically act cooperatively towards a common objective to achieve a team-optimal solution. However, in practice, members may behave selfishly by prioritizing their goals, resulting in an NE under a non-cooperative game. To study this misalignment, we develop a parameterized model for team and game problems, where game parameters represent each individual's deviation from the team objective. The study begins by exploring the consistency and deviation between the NE and the team-optimal solution under fixed game parameters. We provide a necessary and sufficient condition for any NE to be a team optimum, along with establishing an upper bound to measure their difference when this consistency fails. We then study how to steer the NE toward the team-optimal solution by adjusting game parameters in an incomplete-information leader--follower setting, where the leader observes only equilibrium responses rather than the exact lower-level game structure. To address this challenge, we develop a two-stage learned-response intervention framework: the leader first learns a behaviorally consistent lower-level model from observed equilibria and then computes the intervention over the learned response map using bilevel hypergradient optimization, followed by convergence analysis and simulation validation.

math.OC