Search arXiv⌕ Search

arXiv · 2307.11134

Exact convergence rate of the last iterate in subgradient methods

Abstract

We study the convergence of the last iterate in subgradient methods applied to the minimization of a nonsmooth convex function with bounded subgradients. We first introduce a proof technique that generalizes the standard analysis of subgradient methods. It is based on tracking the distance between the current iterate and a different reference point at each iteration. Using this technique, we obtain the exact worst-case convergence rate for the objective accuracy of the last iterate of the projected subgradient method with either constant step sizes or constant step lengths. Tightness is shown with a worst-case instance matching the established convergence rate. We also derive the value of the optimal constant step size when performing $N$ iterations, for which we find that the last iterate accuracy is smaller than $B R \sqrt{1+\log(N)/4}/{\sqrt{N+1}}$ %$\frac{B R \log N}{\sqrt{N+1}}$ , where $B$ is a bound on the subgradient norm and $R$ is a bound on the distance between the initial iterate and a minimizer. Finally, we introduce a new optimal subgradient method that achieves the best possible last-iterate accuracy after a given number $N$ of iterations. Its convergence rate ${B R}/{\sqrt{N+1}}$ matches exactly the lower bound on the performance of any black-box method on the considered problem class. We also show that there is no universal sequence of step sizes that simultaneously achieves this optimal rate at each iteration, meaning that the dependence of the step size sequence in $N$ is unavoidable.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Moslem Zamani, François Glineur. 2023-07-20. Exact convergence rate of the last iterate in subgradient methods. https://arxiv.org/abs/2307.11134

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Symmetry-dependence in Rounding of a Convex Body

The symmetry measure of a convex body $S\subset\mathbb{R}^n$ is given by: $\mathrm{sym}(S):=\max\{α\ge0:\text{ there exists }x\in S\text{ such that }-α(S-x)\subseteq S-x\}$, where such an $x$ is called a Minkowski center. We prove that every convex body $S$ admits a $\sqrt{\frac{n}{\mathrm{sym}(S)}}$-rounding of $S$, namely, there exists an origin-centered ellipsoid $E$ and a center $c$ such that $E\subseteq S-c\subseteq\sqrt{\frac{n}{\mathrm{sym}(S)}}\,E$. This result was conjectured in 2005 by Belloni and Freund. As special cases, this recovers an $n$-rounding of $S$ (since $\mathrm{sym}(S)\ge\frac{1}{n}$), and a $\sqrt{n}$-rounding when $\mathrm{sym}(S)=1$. In the case when $S$ is a polytope given as the convex hull of points, the desired rounding is produced by a regularized minimum-volume covering ellipsoid problem where the regularization is with respect to the Minkowski center. Similarly, when $S$ is a polytope given as the intersection of halfspaces, such a rounding is produced by a regularized maximum-volume inscribed ellipsoid problem. In both of these cases, the rounding can be computed by first solving a linear optimization problem (to compute $\mathrm{sym}(S)$ and a Minkowski center), and then solving a convex optimization problem with a logarithmic determinant objective, second-order cone constraints, and one semidefinite cone constraint. We also show that the factor $\sqrt{\frac{n}{\mathrm{sym}(S)}}$ is nearly tight in its dependence on dimension and symmetry. When $\frac{n+1}{1+\mathrm{sym}(S)}$ is an integer, we show by explicit construction that the factor $\sqrt{\frac{n}{\mathrm{sym}(S)}}$ is tight. In the more general case, for every dimension $n$ and every admissible symmetry value, we construct a polytope $S$ for which every rounding factor is at least $\sqrt{\frac{2}{3}}\sqrt{\frac{n}{\mathrm{sym}(S)}}$.

math.OC↗

Maximal Monotone Differential Inclusions with Volterra and One-Sided Lipschitz Perturbations under Nonlocal Initial Conditions and Applications

We investigate a class of differential inclusions governed by non-autonomous and autonomous maximal monotone operators, involving set-valued perturbations with a Volterra integral term and subject to a nonlocal condition. Under suitable assumptions on the governing operator, the set-valued perturbation, and the Volterra kernel, we establish existence results for solutions. In particular, the set-valued perturbation is assumed to satisfy a one-sided Lipschitz condition, while the Volterra kernel is required to be Lipschitz continuous. The existence of a trajectory is established by means of an iterative construction and an application of Zorn's lemma. Finally, several examples are presented to illustrate the applicability of the abstract results.

math.OC↗

On Fast-Slow Mean-Field Forward-Backward Stochastic Systems

We establish an averaging principle for a class of multiscale mean-field forward-backward stochastic differential equations and identify several novel phenomena that are absent from classical fast-slow systems. In contrast with classical fast-slow systems, the effective dynamics cannot in general be obtained by simply freezing deterministic slow parameters and averaging against the invariant measure of the resulting fast equation. The appropriate averaging object is instead provided by a frozen fast dynamics in a random environment and its associated conditional invariant measures, which retain the coupling between the slow state and its distribution. The forward-backward structure creates a further obstruction: local averaging estimates need not remain stable when propagated over an arbitrary time horizon. We identify a uniform restart stability condition for the averaged system under which this obstruction can be overcome. Using a joint lifted semigroup for the state-law dynamics, together with a two-scale discretization and a Gordin-type decomposition, we prove strong averaging for both the forward and backward components with optimal convergence rate $O(\varepsilon^{1/2})$. As an application, we apply the general theory to a class of mean-field stochastic control problems and develop an efficient algorithm for solving such mean-field control problems.

math.OC↗