Search arXiv⌕ Search

arXiv · 2610.01662

Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex--Concave Minimax Optimization

Abstract

We establish complexity lower bounds for stochastic first-order algorithms in nonconvex--concave minimax optimization, allowing algorithms to use variance reduction. Our main contribution is a lower bound for a zero-respecting algorithm class that permits variance reduction, extending beyond the algorithmic restrictions imposed by some existing lower bounds. We consider objectives with an $L$-Lipschitz continuous joint gradient, a compact convex dual domain of Euclidean radius at most $D_Y$, and a primal value function, defined by maximizing the objective over the dual variable, with initial suboptimality at most $Δ$. The target accuracy $\varepsilon$ is measured by the gradient norm of the Moreau envelope of the constrained primal value function with parameter $1/(2L)$. Under an unbiased stochastic first-order oracle with variance at most $σ^2$ and mean-square smoothness, we prove the lower bound $Ω\!\left(L^2D_YΔ\varepsilon^{-3}+L^3D_Y^2Δσ^2\varepsilon^{-6}\right)$. This result quantifies the dependence on accuracy, dual-domain radius, and oracle noise even when variance reduction is allowed. We also establish complementary lower bounds for nonconvex--strongly-concave minimax optimization. With dual strong-concavity parameter $μ>0$ and condition number $κ:=L/μ$, we obtain $Ω\!\left(LΔ\sqrtκ\,\varepsilon^{-2}+LΔκσ^2\varepsilon^{-4}\right)$ under the bounded-variance oracle model. Under the additional mean-square smoothness condition with constant $\bar L$, we obtain $Ω\!\left(LΔ\sqrtκ\,\varepsilon^{-2}+Δ\bar Lσκ^{3/2}\varepsilon^{-3}\right)$. Together, these results identify complexity barriers across the concave and strongly concave regimes, with the main nonconvex--concave bound remaining valid for algorithms that use variance reduction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jiayi Song, Zi Xu. 2026-10-01. Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex--Concave Minimax Optimization. https://arxiv.org/abs/2610.01662

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Using Less for More: When Warm-Starting Accelerates Branch-and-Cut for Stochastic Programs

Two-stage stochastic programs quickly become intractable as the number of scenarios grows. Motivated by this, we propose TULIP, a modular and easy-to-implement three-step warm-start framework for two-stage stochastic (mixed-)integer programs with an exponential number of cuts separated during branch-and-cut. TULIP (a) builds a cheap surrogate of the full problem by reducing the scenario set or by decoupling the two stages, (b) solves it up to a first incumbent to collect the tight cuts separated along the way, and (c) injects them to warm-start the original problem. In short: we use less (a cheaper surrogate) for more (the original problem). Using this modular setup, we propose four methods within this framework, each with a slightly different setting. Across four case studies, we show that this acceleration is governed by a single mechanism, the root cut loop, and we specify it through a closed-form equation. This TULIP speedup model predicts a speedup when the time saved in the root cut loop exceeds the surrogate overhead. In our experiments, a TULIP variant achieves mean speedups of up to 2.56, with gains increasing with the scenario count. In the remaining case studies, TULIP provides little or no runtime benefit, which the TULIP speedup model mostly explains through insufficient root cut loop savings compared to the surrogate overhead.

math.OC↗

Convergence Analysis of the Wasserstein Proximal Algorithm beyond Geodesic Convexity

The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent progress in studying the training dynamics of the noisy gradient descent algorithm on two-layer neural networks in the mean-field regime, we provide in this paper a simple and self-contained analysis for the convergence of the general-purpose Wasserstein proximal algorithm without assuming geodesic convexity of the objective functional. Under a natural Wasserstein analog of the Euclidean Polyak-Łojasiewicz inequality, we establish that the proximal algorithm achieves an unbiased and linear convergence rate. Our convergence rate improves upon existing rates of the proximal algorithm for solving Wasserstein gradient flows under strong geodesic convexity. We also extend our analysis to the inexact proximal algorithm for geodesically semiconvex objectives. In our numerical experiments, proximal training demonstrates a faster convergence rate than the noisy gradient descent algorithm on mean-field neural networks.

math.OC↗

Technological foundations of management decision-making in the reconstruction of complex gas pipeline system

This monograph presents a comprehensive analysis of the technological foundations of management decision-making in the reconstruction of complex gas pipeline systems. The study addresses the challenges posed by the aging infrastructure of gas supply networks and explores advanced strategies to improve their reliability, efficiency, and automation. Particular attention is given to the reconstruction of pipelines with various configurations linear, looped, and parallel systems under non-stationary gas flow conditions. The proposed models and methodologies offer solutions for optimizing operational parameters, improving emergency valve response, and ensuring uninterrupted gas supply through advanced management systems and data-driven decision support tools. Emphasis is placed on the integration of modern technologies, system theory, and feedback mechanisms in the design and operation of reconstructed pipeline systems. This work is intended for engineers, system designers, and researchers in the fields of gas supply, systems engineering, and energy infrastructure.

math.OC↗