Search arXivSearch

arXiv · 2101.09545

Acceleration Methods

Abstract

This monograph covers some recent advances in a range of acceleration techniques frequently used in convex optimization. We first use quadratic optimization problems to introduce two key families of methods, namely momentum and nested optimization schemes. They coincide in the quadratic case to form the Chebyshev method. We discuss momentum methods in detail, starting with the seminal work of Nesterov and structure convergence proofs using a few master templates, such as that for optimized gradient methods, which provide the key benefit of showing how momentum methods optimize convergence guarantees. We further cover proximal acceleration, at the heart of the Catalyst and Accelerated Hybrid Proximal Extragradient frameworks, using similar algorithmic patterns. Common acceleration techniques rely directly on the knowledge of some of the regularity parameters in the problem at hand. We conclude by discussing restart schemes, a set of simple techniques for reaching nearly optimal convergence rates while adapting to unobserved regularity parameters.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alexandre d'Aspremont, Damien Scieur, Adrien Taylor. 2024-09-24. Acceleration Methods. https://doi.org/10.1561/2400000036

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stochastic Optimization Algorithms for Problems with Controllable Biased Oracles

Motivated by emerging applications in machine learning, we consider an optimization problem in a general setting in which the gradient of the objective function is available via a biased stochastic oracle. We assume a bias-control parameter can reduce the bias magnitude; however, a lower bias requires more computation/samples. For instance, in two applications on stochastic composition optimization and policy optimization for infinite-horizon Markov decision processes, we show that the bias follows a power law and exponential decay, respectively, as functions of their corresponding bias control parameters. For problems with such gradient oracles, the paper proposes stochastic algorithms that adjust the bias-control parameter throughout the iterations. We analyze the nonasymptotic performance of the proposed algorithms in the nonconvex regime and establish their sample or bias-control computation complexities to obtain a stationary point in expectation or with high probability. Finally, we numerically evaluate the performance of the proposed algorithms over three applications.

math.OC

Inverse Problems Over Probability Measure Space

Define a forward problem as $ρ_y = G_\#ρ_x$, where the probability distribution $ρ_x$ is mapped to another distribution $ρ_y$ using the forward operator $G$. In this work, we investigate the corresponding inverse problem: Given $ρ_y$, how to find $ρ_x$? Depending on whether $ G$ is overdetermined or underdetermined, the solution can have drastically different behavior. In the overdetermined case, we formulate a variational problem $\min_{ρ_x} D( G_\#ρ_x, ρ_y)$, and find that different choices of the metric $ D$ significantly affect the quality of the reconstruction. When $ D$ is set to be the Wasserstein distance, the reconstruction is the marginal distribution, while setting $ D$ to be a $ϕ$-divergence reconstructs the conditional distribution. In the underdetermined case, we formulate the constrained optimization $\min_{\{ G_\#ρ_x=ρ_y\}} E[ρ_x]$. The choice of $ E$ also significantly impacts the construction: setting $ E$ to be the entropy gives us the piecewise constant reconstruction, while setting $ E$ to be the second moment, we recover the classical least-norm solution. We also examine the formulation with regularization: $\min_{ρ_x} D( G_\#ρ_x, ρ_y) + α\mathsf R[ρ_x]$, and find that the entropy-entropy pair leads to a regularized solution that is defined in a piecewise manner, whereas the $W_2$-$W_2$ pair leads to a least-norm solution where $W_2$ is the 2-Wasserstein metric.

math.OC

When Does Selfishness Align with Team Goals? A Structural Analysis of Equilibrium and Optimality

This paper investigates the relationship between the team-optimal solution and the Nash equilibrium (NE) to assess the impact of self-interested decisions on team performance. In classical team decision problems, team members typically act cooperatively towards a common objective to achieve a team-optimal solution. However, in practice, members may behave selfishly by prioritizing their goals, resulting in an NE under a non-cooperative game. To study this misalignment, we develop a parameterized model for team and game problems, where game parameters represent each individual's deviation from the team objective. The study begins by exploring the consistency and deviation between the NE and the team-optimal solution under fixed game parameters. We provide a necessary and sufficient condition for any NE to be a team optimum, along with establishing an upper bound to measure their difference when this consistency fails. We then study how to steer the NE toward the team-optimal solution by adjusting game parameters in an incomplete-information leader--follower setting, where the leader observes only equilibrium responses rather than the exact lower-level game structure. To address this challenge, we develop a two-stage learned-response intervention framework: the leader first learns a behaviorally consistent lower-level model from observed equilibria and then computes the intervention over the learned response map using bilevel hypergradient optimization, followed by convergence analysis and simulation validation.

math.OC