Search arXivSearch

arXiv · 2004.00475

Stopping Criteria for, and Strong Convergence of, Stochastic Gradient Descent on Bottou-Curtis-Nocedal Functions

Abstract

Stopping criteria for Stochastic Gradient Descent (SGD) methods play important roles from enabling adaptive step size schemes to providing rigor for downstream analyses such as asymptotic inference. Unfortunately, current stopping criteria for SGD methods are often heuristics that rely on asymptotic normality results or convergence to stationary distributions, which may fail to exist for nonconvex functions and, thereby, limit the applicability of such stopping criteria. To address this issue, in this work, we rigorously develop two stopping criteria for SGD that can be applied to a broad class of nonconvex functions, which we term Bottou-Curtis-Nocedal functions. Moreover, as a prerequisite for developing these stopping criteria, we prove that the gradient function evaluated at SGD's iterates converges strongly to zero for Bottou-Curtis-Nocedal functions, which addresses an open question in the SGD literature. As a result of our work, our rigorously developed stopping criteria can be used to develop new adaptive step size schemes or bolster other downstream analyses for nonconvex functions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vivak Patel. 2021-04-01. Stopping Criteria for, and Strong Convergence of, Stochastic Gradient Descent on Bottou-Curtis-Nocedal Functions. https://arxiv.org/abs/2004.00475

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Nonsmooth Newton methods with effective subspaces for polyhedral regularization

We propose several new nonsmooth Newton methods for solving convex composite optimization problems with polyhedral regularizers, while avoiding the computation of complicated second-order information on these functions. Under the tilt-stability condition at the optimal solution, these methods achieve the quadratic convergence rates expected of Newton schemes. Numerical experiments on Lasso, generalized Lasso, OSCAR-regularized least-square problems, and an image super-resolution task illustrate both the broad applicability and the accelerated convergence profile of the proposed algorithms, in comparison with first-order and several recently developed nonsmooth Newton schemes.

math.OC

Non-Asymptotic Global Convergence of PPO-Clip

Reinforcement learning has gained attention for modern Large Language Model post-training. The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their efficiency. These algorithms incorporate a clipping mechanism to improve stability. Besides, a regularization term, such as the reverse KL-divergence or a more general \(f\)-divergence, is introduced to control excessive deviation from a reference policy. Despite their empirical success, a rigorous theoretical understanding of the problem and the algorithm's properties is limited. This paper advances the theoretical foundations of the PPO-Clip algorithm by analyzing a deterministic actor-only PPO algorithm within the general RL setting with \(f\)-divergence regularization under the softmax policy parameterization. We derive a non-uniform Lipschitz smoothness condition and a Łojasiewicz inequality for the considered problem. Based on these properties, we establish non-asymptotic global linear convergence in value gap for the forward KL regularizer. For the reverse KL regularizer, we derive global linear convergence from any finite softmax initialization in both value gap and squared policy distance.

math.OC

An Occupation-Measure and Frank-Wolfe Framework for Heterogeneous Mean-Field Control

Heterogeneous mean-field control (MFC) involves multiple interacting populations with distinct dynamics, control constraints, and coupling structures. We develop a heterogeneous occupation-measure mean-field control (OM-MFC) framework that reformulates the population-distribution control problem as an optimization problem over occupation measures. In this representation, potentially nonlinear dynamics enter through linear weak Liouville constraints, while nonlocal inter-population interactions become structured bilinear functionals of the occupation-measure marginals. Exploiting this structure, we derive a Frank--Wolfe (FW) method and show that its linear minimization oracle decomposes into independent population-wise optimal control problems once the current population distributions are fixed, enabling parallel computation without an a priori discretization of the measure space. We further establish a sufficient convexity condition based on the symmetrized matrix-valued interaction kernel, under which Frank--Wolfe with an exact oracle admits the standard $O(1/k)$ convergence rate. Numerical examples on heterogeneous UAV coordination and three-dimensional search-and-rescue illustrate the framework both within the certified convex regime and for directional interactions beyond that regime.

math.OC