Search arXiv⌕ Search

arXiv · 2609.34099

Constrained Nonconvex Stochastic Optimization with One Projection

Abstract

Constrained nonconvex optimization has seen increasing application in modern machine learning, such as safe LLM alignment/finetuning, transfer learning and rank-constrained continual learning. The commonly used approach is Projected SGD which requires projection at every iteration. However, projection onto a functional constraint can cost substantially more than a stochastic first-order update. We study whether this operation can be deferred until the end of nonconvex stochastic optimization, and only do it once. For weakly convex, possibly nonsmooth objectives with regular convex constraints, we give a penalized proximal method that uses $\widetilde O(ε^{-4})$ stochastic subgradients and one projection onto the constraint set. The returned point is exactly feasible and has a small Moreau-envelope stationarity measure; for smooth objectives, the same algorithm controls the projected-gradient mapping. For smooth nonconvex constraints, we assume a global lower bound on the infeasible constraint slope which permits arbitrary, possibly infeasible initialization. An exact-penalty variant attains the same stochastic-oracle order and returns an exactly feasible point within $O(ε)$ of an $ε$-KKT point. All intermediate updates use only projections onto a simple Euclidean ball. These are oracle-complexity guarantees: the cost of the single terminal projection is separate, and no efficient projection algorithm for a general nonconvex set is assumed to follow from regularity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuyang Deng, Mohammadreza M. Kalan, Eitan J. Neugut, Mehrdad Mahdavi. 2026-09-28. Constrained Nonconvex Stochastic Optimization with One Projection. https://arxiv.org/abs/2609.34099

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adaptive Extremum Seeking Control via the RMSprop Optimizer

Extremum Seeking Control (ESC) is a well-known set of continuous time algorithms for model-free optimization of a cost function. One issue for ESCs is the convergence rates of parameters to extrema of unknown cost functions. The local convergence rate depends on the second, or sometimes higher, order derivatives of the unknown cost function. To mitigate this dependency, we propose the use of the RMSprop optimizer for ESCs as RMSprop is an adaptive gradient-based optimizer which attempts to have a normalized convergence rate in all parameters. Practical stability results are given for this RMSprop ESC (RMSpESC). In particular notability, the proof of practical stability uses Lyapunov function based on observed contracting, attractive sets. Versions of this Lyapunov function could be applied to other areas of applications, in particular for interconnected systems.

math.OC↗

A deterministic optimization algorithm for nonconvex and combinatorial bi-objective programming

Many practical multiobjective optimization (MOO) problems include discrete decision variables and/or nonlinear model equations and exhibit disconnected or smooth but nonconvex Pareto surfaces. Scalarization methods, such as the weighted-sum and sandwich (SD) algorithms, are common approaches to solving MOO problems but may fail on nonconvex or discontinuous Pareto fronts. In the current work, motivated by the well-known normal boundary intersection (NBI) method and the SD algorithm, we present SDNBI, a new algorithm for bi-objective optimization (BOO) designed to address the theoretical and numerical challenges associated with the reliable solution of general nonconvex and discrete BOO problems. The main improvements in the algorithm are the effective exploration of the nonconvex regions of the Pareto front and, uniquely, the early identification of regions where no additional Pareto solutions exist. The performance of the SDNBI algorithm is assessed based on the accuracy of the approximation of the Pareto front constructed over the disconnected nonconvex objective domains. The new algorithm is compared with two MOO approaches, the modified NBI method and the SD algorithm, using published benchmark problems. The results indicate that the SDNBI algorithm outperforms the modified NBI and SD algorithms in solving convex, nonconvex-continuous, and combinatorial problems, both in terms of computational cost and of the overall quality of the Pareto-optimal set, suggesting that the SDNBI algorithm is a promising alternative for solving BOO problems.

math.OC↗

ICNN-enhanced 2SP: Leveraging input convex neural networks for solving two-stage stochastic programming

Two-stage stochastic programming (2SP) offers a basic framework for modelling decision-making under uncertainty, yet scalability remains a challenge due to the computational complexity of recourse function evaluation. Existing learning-based methods like Neural Two-Stage Stochastic Programming (Neur2SP) employ neural networks (NNs) as recourse function surrogates but rely on computationally intensive mixed-integer programming (MIP) formulations. We propose ICNN-enhanced 2SP, a method that leverages Input Convex Neural Networks (ICNNs) to exploit linear programming (LP) representability in convex 2SP problems. By architecturally enforcing convexity and enabling exact inference through LP, our approach eliminates the need for integer variables inherent in the conventional MIP-based formulation while retaining an exact embedding of the ICNN surrogate within the 2SP framework. This results in a more computationally efficient alternative, and we show that good solution quality can be maintained. Comprehensive experiments reveal that ICNNs incur only marginally longer training times while achieving validation accuracy on par with their standard NN counterparts. Across benchmark problems, ICNN-enhanced 2SP often exhibits considerably faster solution times than the MIP-based formulations while preserving solution quality, with these advantages becoming significantly more pronounced as problem scale increases. For the most challenging instances, the method achieves speedups of up to 100$\times$ with solution quality superior to MIP-based formulations.

math.OC↗