Search arXivSearch

arXiv · 2307.15868

Faster Stochastic Algorithms for Minimax Optimization under Polyak--Łojasiewicz Conditions

Abstract

This paper considers stochastic first-order algorithms for minimax optimization under Polyak--Łojasiewicz (PL) conditions. We propose SPIDER-GDA for solving the finite-sum problem of the form $\min_x \max_y f(x,y)\triangleq \frac{1}{n} \sum_{i=1}^n f_i(x,y)$, where the objective function $f(x,y)$ is $μ_x$-PL in $x$ and $μ_y$-PL in $y$; and each $f_i(x,y)$ is $L$-smooth. We prove SPIDER-GDA could find an $ε$-optimal solution within ${\mathcal O}\left((n + \sqrt{n}\,κ_xκ_y^2)\log (1/ε)\right)$ stochastic first-order oracle (SFO) complexity, which is better than the state-of-the-art method whose SFO upper bound is ${\mathcal O}\big((n + n^{2/3}κ_xκ_y^2)\log (1/ε)\big)$, where $κ_x\triangleq L/μ_x$ and $κ_y\triangleq L/μ_y$. For the ill-conditioned case, we provide an accelerated algorithm to reduce the computational cost further. It achieves $\tilde{\mathcal O}\big((n+\sqrt{n}\,κ_xκ_y)\log (κ_y/ε) \log(1/ε)\big)$ SFO upper bound when $κ_y \gtrsim \sqrt{n}$. Our ideas can also be applied to a more general setting where the objective function only satisfies the PL condition for one variable. Numerical experiments validate the superiority of proposed methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lesi Chen, Boyuan Yao, Luo Luo. 2026-03-15. Faster Stochastic Algorithms for Minimax Optimization under Polyak--Łojasiewicz Conditions. https://arxiv.org/abs/2307.15868

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AdGT: Decentralized Gradient Tracking with Adaptive Per-Agent Stepsizes

In decentralized optimization, gradient-tracking methods typically rely on a single global stepsize. This choice can be conservative when agents have local objectives with different smoothness constants, since the stepsize must remain stable for the agent with the largest smoothness constant. This paper proposes AdGT, a decentralized gradient-tracking method in which each agent adapts its own stepsize using local gradient variation and a single global safety factor. The method reduces fixed-stepsize tuning effort and allows agents to exploit local smoothness information during the iterations. For smooth and strongly convex local objectives over undirected networks, we prove that the analyzed AdGT update converges linearly to the exact consensus optimizer. We also study two adaptive stepsize updates that use changes in the gradient-tracking direction. We characterize when the corresponding candidate determines the stepsize and prove conditional lower and upper stepsize bounds and linear convergence under an additional relative tracking-disagreement condition. Experiments on logistic regression, ridge regression, synthetic quadratic problems, and a linear-regression benchmark against state-of-the-art decentralized solvers show that AdGT often reaches a given accuracy in fewer iterations or gradient evaluations than tuned fixed-stepsize GT and the tested baselines, especially under heterogeneous local smoothness. In the topology experiments, each tested AdGT update uses one common safety factor across all graphs, whereas the fixed GT stepsize is tuned separately for each graph and seed.

math.OC

Brockett cost function for symplectic eigenvalues

The sum of symplectic eigenvalues and corresponding eigenvectors of symmetric positive-definite matrices in the sense of Williamson's theorem can be computed via minimization of a trace cost function under the symplecticity constraint. Optimal solutions to this problem only offer a symplectic basis for the symplectic eigenspace corresponding to the sought symplectic eigenvalues. In this note, we introduce a Brockett cost function and investigate its properties and the connection with the symplectic eigenvalues and eigenvectors of the considered matrix. Specifically, we prove that any stationary point consists of symplectic eigenvectors, characterize the saddle points and global minimizers based on which the trace minimization theorem for the symplectic eigenvalues is re-established, and the nonexistence of local nonglobal minimizers is justified.

math.OC

Riemannian Bilevel Optimization with Gradient Aggregation

We study bilevel optimization on Riemannian manifolds when the lower-level solution set is a positive-dimensional submanifold, so that implicit differentiation fails. We propose Riemannian Bilevel Descent Aggregation (RBDA), which extends bilevel descent aggregation to manifolds. Its inner loop aggregates the lower-level descent direction with the upper-level gradient under a decaying multiplier, and its hypergradient is the reverse-mode derivative of the unrolled loop. Under geodesic convexity and quadratic growth of the lower level, the inner iterates converge to a point of the optimistic solution set at a polynomial rate. Approximate minimizers of the objective with a finite number of inner iterations converge to minimizers of the optimistic value. In the experiments RBDA selects the optimistic solution where the implicit and unrolled estimators remain at the initial point or stop at a larger query loss. While each of its outer steps costs more than that of the unrolled estimator, it attains the highest test accuracy in data hyper-cleaning.

math.OC