Search arXivSearch

arXiv · 2005.05645

Convergence of Online Adaptive and Recurrent Optimization Algorithms

Abstract

We prove local convergence of several notable gradient descent algorithms used in machine learning, for which standard stochastic gradient descent theory does not apply directly. This includes, first, online algorithms for recurrent models and dynamical systems, such as \emph{Real-time recurrent learning} (RTRL) and its computationally lighter approximations NoBackTrack and UORO; second, several adaptive algorithms such as RMSProp, online natural gradient, and Adam with $β^2\to 1$.Despite local convergence being a relatively weak requirement for a new optimization algorithm, no local analysis was available for these algorithms, as far as we knew. Analysis of these algorithms does not immediately follow from standard stochastic gradient (SGD) theory. In fact, Adam has been proved to lack local convergence in some simple situations \citep{j.2018on}. For recurrent models, online algorithms modify the parameter while the model is running, which further complicates the analysis with respect to simple SGD.Local convergence for these various algorithms results from a single, more general set of assumptions, in the setup of learning dynamical systems online. Thus, these results can cover other variants of the algorithms considered.We adopt an "ergodic" rather than probabilistic viewpoint, working with empirical time averages instead of probability distributions. This is more data-agnostic and creates differences with respect to standard SGD theory, especially for the range of possible learning rates. For instance, with cycling or per-epoch reshuffling over a finite dataset instead of pure i.i.d.\ sampling with replacement, empirical averages of gradients converge at rate $1/T$ instead of $1/\sqrt{T}$ (cycling acts as a variance reduction method), theoretically allowing for larger learning rates than in SGD.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pierre-Yves Massé, Yann Ollivier. 2021-01-08. Convergence of Online Adaptive and Recurrent Optimization Algorithms. https://arxiv.org/abs/2005.05645

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Equidistribution of saddle periodic points for Hénon-like maps

We prove that under a natural assumption on the dynamical degrees, the saddle periodic points of a Hénon-like map in any dimension equidistribute with respect to the equilibrium measure. Our work is a generalization of the results of Bedford-Lyubich-Smillie, Dujardin, and Dinh-Sibony along with improvements of their techniques. We also investigate some fine properties of Green currents associated with the map.

math.DS

On dissonance and orthogonal projections of self-conformal measures

Let $μ$ be a self-conformal measure on $\mathbb{R}^d$. We establish conditions for $μ$ under which $\dim(μ*ν) = \min\lbrace d,\dimμ+\dimν\rbrace$ holds when $ν$ is any Ahlfors-regular or self-conformal measure on $\mathbb{R}^d$. Our main result states the following sufficient condition: $μ$ is totally non-linear and not supported on a smooth hypersurface. We also establish sufficient (likely non-sharp) algebraic conditions for self-conformal measures which are not totally non-linear. In addition, we show that $\dim μ\circπ^{-1} = \min\{ k, \dim μ\}$ for every ortohogonal projection $π:\mathbb{R}^d\to\mathbb{R}^k$, $0<k<d$, when either $d=2$ and $μ$ is not self-similar and not supported on a line, or $d\geq 3$ and $μ$ is totally non-linear and not supported on a smooth hypersurface.

math.DS

Equation-Free Screening of Mittag-Leffler-Compatible Dynamics from Scalar Time Series via kNN Multi-Horizon Profiles

Fractional models provide a natural description of systems with memory, but a noninteger derivative should not be introduced solely because a time series is curved or slowly relaxing. We develop an equation-free preliminary screening framework that asks whether a scalar time series produces a multi-horizon k-nearest-neighbor (kNN) profile more compatible with Mittag-Leffler-type behavior than with selected conventional alternatives. In an ideal matched Caputo-relaxation benchmark, the complete generation-kNN-profile-model-comparison pipeline reproduces the expected Mittag-Leffler geometry and recovers the generating order to within approximately $10^{-3}$; this is interpreted as controlled calibration rather than as general fractional-order identification. Under 3% trajectory-specific observational noise, the held-out Mittag-Leffler preference is most consistent when the generating dynamics are well separated from the integer-order limit and becomes progressively less decisive as $α\rightarrow1$. The fitted order $α_{\mathrm{fit}}$, however, shows substantially larger realization-to-realization variability. Thus, relative model compatibility is more robust than single-realization order estimation in the present noisy benchmark. Noise-free nonfractional controls show a separate limitation of specificity: a stretched exponential can generate a strongly Mittag-Leffler-compatible profile, whereas inclusion of the generating rational/Hill family recovers that family and its parameters to numerical precision in the matched setting. A positive Mittag-Leffler-versus-exponential screen therefore does not uniquely establish fractional origin. A fractional chaotic system is treated only as an exploratory extension: the Mittag-Leffler growth family gives lower finite-window RMSE than exponential and logistic/saturating alternatives over the detected pre-transition interval.

math.DS