Search arXiv⌕ Search

arXiv · 2609.36600

Second-Moment Stochastic Approximation Methods

Abstract

Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky's theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tao Jiang, Lin Xiao. 2026-09-29. Second-Moment Stochastic Approximation Methods. https://arxiv.org/abs/2609.36600

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Douglas--Rachford for multioperator comonotone inclusions with applications to multiblock optimization

We study the doubly relaxed Douglas--Rachford (DR) algorithm for solving a multioperator inclusion problem involving the sum of maximally comonotone operators. To address such problems, we adopt a product space reformulation that accommodates nonconvex-valued operators, which is essential when dealing with weakly comonotone mappings. We establish the convergence of the doubly relaxed DR algorithm under comonotonicity assumptions, subject to suitable conditions on the algorithm parameters and the comonotonicity moduli of the operators. Our analysis relies on the Attouch--Théra duality framework, which enables the study of convergence through the corresponding dual inclusion problem. As an application, we derive a multiblock ADMM-type algorithm for structured convex and nonconvex optimization problems by applying the doubly relaxed DR algorithm to the operator inclusion formulation of the KKT system. The resulting method extends the classical duality between the DR algorithm and the alternating direction method of multipliers from the convex two-block case to multiblock and nonconvex settings. Moreover, we establish convergence guarantees in both the fully convex and strongly convex-weakly convex regimes.

math.OC↗

Decentralized Optimization over Time-Varying Row-Stochastic Digraphs

Decentralized optimization over directed graphs underlies applications such asrobotic swarms, sensor networks, and distributed learning. In many such systems, the network is a Time-Varying Broadcast Network (TVBN), in which out-degrees are unknown and only row-stochastic mixing matrices can beconstructed. Exact convergence of decentralized optimization over TVBNs has remained a long-standing open problem. Row-stochastic mixing converges to aweighted average given by the limit vector of the matrix product; since this vector depends on unpredictable future graph realizations, bias-correction techniques that estimate it are infeasible. We develop the first decentralized optimization algorithm that converges exactly using only time-varying row-stochastic matrices. Its core is PULM (Pull-with-Memory), a gossip protocol based on a different principle: a limit vector that is not yet determined can be controlled rather than estimated. PULM interleaves row-stochastic gossip with a communication-free adjustment in which each of the $n$ nodes anchors the weight of its initial vector at $1/n$, achieving exponentially fast average consensus for every admissible graph sequence. Building on PULM, PULM-DGD finds a solution with squared gradient norm at most $ε$ for smooth nonconvex objectives within $\mathcal{O}(ε^{-1}\ln(1/ε))$ communication rounds, extending decentralized optimization to highly dynamic networks.

math.OC↗

Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization

We study decentralized online optimization of upper-linearizable payoffs over an action set under efficient separation access, with applications to online continuous diminishing-return (DR) submodular maximization. We propose Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evaluate each agent's played action against the average of all local objectives. Each agent maps an internal iterate to a feasible action through an approximate gauge projection, communicates only a cumulative surrogate-gradient dual state, and invokes the local HybridNewton procedure to approximately minimize its post-communication BFTRL potential. For every agent, we achieve expected network-aggregate regret of $\widetilde O(\sqrt{T})$. Over $T$ rounds, each agent uses $T$ neighbor-mixing steps and $\widetilde O(T)$ separation-oracle calls. We give wrapper instantiations covering four up-concave or DR-submodular maximization problems.

math.OC↗