Search arXivSearch

arXiv · 2309.02944

The case for and against fixed step-size: Stochastic approximation algorithms in optimization and machine learning

Abstract

Theory and application of stochastic approximation (SA) have become increasingly relevant due in part to applications in optimization and reinforcement learning. This paper takes a new look at SA with constant step-size $α>0$, defined by the recursion, $$θ_{n+1} = θ_{n}+ αf(θ_n,Φ_{n+1})$$ in which $θ_n\in\mathbb{R}^d$ and $\{Φ_{n}\}$ is a Markov chain. The goal is to approximately solve root finding problem $\bar{f}(θ^*)=0$, where $\bar{f}(θ)=\mathbb{E}[f(θ,Φ)]$ and $Φ$ has the steady-state distribution of $\{Φ_{n}\}$. The following conclusions are obtained under an ergodicity assumption on the Markov chain, compatible assumptions on $f$, and for $α>0$ sufficiently small: $\textbf{1.}$ The pair process $\{(θ_n,Φ_n)\}$ is geometrically ergodic in a topological sense. $\textbf{2.}$ For every $1\le p\le 4$, there is a constant $b_p$ such that $\limsup_{n\to\infty}\mathbb{E}[\|θ_n-θ^*\|^p]\le b_p α^{p/2}$ for each initial condition. $\textbf{3.}$ The Polyak-Ruppert-style averaged estimates $θ^{\text{PR}}_n=n^{-1}\sum_{k=1}^{n}θ_k$ converge to a limit $θ^{\text{PR}}_\infty$ almost surely and in mean square, which satisfies $θ^{\text{PR}}_\infty=θ^*+α\barΥ^*+O(α^2)$ for an identified non-random $\barΥ^*\in\mathbb{R}^d$. Moreover, the covariance is approximately optimal: The limiting covariance matrix of $θ^{\text {PR}}_n$ is approximately minimal in a matricial sense. The two main take-aways for practitioners are application-dependent. It is argued that, in applications to optimization, constant gain algorithms may be preferable even when the objective has multiple local minima; while a vanishing gain algorithm is preferable in applications to reinforcement learning due to the presence of bias.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Caio Kalil Lauand, Ioannis Kontoyiannis, Sean Meyn. 2025-11-10. The case for and against fixed step-size: Stochastic approximation algorithms in optimization and machine learning. https://arxiv.org/abs/2309.02944

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The level of self-organized criticality in oscillating Brownian motion: $n$-consistency and stable Poisson-type convergence of the MLE

For some discretely observed path of oscillating Brownian motion with level of self-organized criticality $ρ_0$, we prove in the infill asymptotics that the MLE is $n$-consistent, where $n$ denotes the sample size, and derive its limit distribution with respect to stable convergence. As the transition density of this homogeneous Markov process is not even continuous in $ρ_0$, the analysis is highly non-standard. Therefore, interesting and somewhat unexpected phenomena occur: The likelihood function splits into several components, each of them contributing very differently depending on how close the argument $ρ$ is to $ρ_0$. Correspondingly, the MLE is successively excluded to lay outside a compact set, a $1/\sqrt{n}$-neighborhood and finally a $1/n$-neighborhood of $ρ_0$ asymptotically. The crucial argument to derive the stable convergence is to exploit the semimartingale structure of the sequential suitably rescaled local log-likelihood function (as a process in time). Both sequentially and as a process in $ρ$, it exhibits a bivariate Poissonian behavior in the stable limit with its intensity being a multiple of the local time at $ρ_0$.

math.ST

Statistical models as natural transformations: meaningfulness, coherence and priors as states in Markov categories

We show that a statistical model in the sense of McCullagh, in the form given by Brøns, is a natural transformation between two functors from the category of designs to the Kleisli category Stoch of the Giry monad, provided that its components are measurable in the parameter. The condition is empty for finite models. A design-indexed quantity is a family of morphisms of Stoch defined on the parameter objects, called meaningful if it is natural. We prove that Tjur's criterion, imposed on parameter functions indexed by finite samples with multiplicities, forces the indexing by the support and then coincides with naturality over the insertions. For finite designs we show that a quantity can be corrected to a natural one within a given class of corrections if and only if a class vanishes in the first cohomology group of a Baues-Wirsching complex relative to that class, while its image in the absolute group is always zero. In the one-way layout, marginal dispersion is not meaningful, and within-group dispersion is the unique correction that leaves the merged design unchanged. A prior is a family of states on the parameter objects, called coherent over a class of design morphisms if it is natural over that class. We show that coherence at a merge confines the prior to the image of the corresponding parameter map, that coherence over the insertions is Kolmogorov consistency, and that coherence over the injections adds the exchangeability assumed by the categorical de Finetti theorem. In the finite one-way scheme, the coherent priors form polytopes of known dimension. The analogue of Jeffreys' general rule is not coherent, while the analogue for location-scale families is. Finally, we show that ridge regression is the Bayesian inversion of the Gaussian linear model with respect to a Gaussian prior, which is coherent over the insertions and never over the injections.

math.ST

Inference for H{ü}sler-Reiss block models

Estimating the H{ü}sler-Reiss precision matrix is a fundamental problem for statistical inference in multivariate extremes. In high-dimensional settings, the number of unknown parameters grows quadratically with the dimension, making regularisation indispensable. Existing approaches regularise the estimation problem by exploiting sparsity. In this paper, we consider an alternative structural assumption, namely that the H{ü}sler-Reiss precision matrix is block-structured. To estimate such a matrix, we introduce a new regularisation framework based on a convex fusion penalty. By encouraging rows and columns to merge, this approach provides a parsimonious representation of the precision matrix, allowing for the simultaneous estimation of its coefficients and the underlying partition of the variables. The resulting convex optimisation problem is solved by an efficient algorithm combining gradient-based updates with progressive fusion steps. We establish non-asymptotic concentration bounds for the empirical weights entering the penalty and prove consistency of both block recovery and precision matrix estimation under suitable regularity conditions. Numerical experiments demonstrate that our methodology accurately recovers the latent block structure while accurately estimating the H{ü}sler-Reiss precision matrix across various configurations, illustrating the practical benefits of fusion-based regularisation for multivariate extremes. These benefits are also demonstrated by applying the proposed method to foreign exchange data.

math.ST