Search arXivSearch

arXiv · 2103.12293

Stochastic Reweighted Gradient Descent

Abstract

Despite the strong theoretical guarantees that variance-reduced finite-sum optimization algorithms enjoy, their applicability remains limited to cases where the memory overhead they introduce (SAG/SAGA), or the periodic full gradient computation they require (SVRG/SARAH) are manageable. A promising approach to achieving variance reduction while avoiding these drawbacks is the use of importance sampling instead of control variates. While many such methods have been proposed in the literature, directly proving that they improve the convergence of the resulting optimization algorithm has remained elusive. In this work, we propose an importance-sampling-based algorithm we call SRG (stochastic reweighted gradient). We analyze the convergence of SRG in the strongly-convex case and show that, while it does not recover the linear rate of control variates methods, it provably outperforms SGD. We pay particular attention to the time and memory overhead of our proposed method, and design a specialized red-black tree allowing its efficient implementation. Finally, we present empirical results to support our findings.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ayoub El Hanchi, David A. Stephens. 2021-03-23. Stochastic Reweighted Gradient Descent. https://arxiv.org/abs/2103.12293

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization

This paper introduces the notion of upper-linearizable/quadratizable functions, a class that extends concavity and DR-submodularity in various settings, including monotone and non-monotone cases. A general meta-algorithm is devised to convert algorithms for linear/quadratic maximization into ones that optimize upper-linearizable/quadratizable functions, offering a unified approach to tackling concave and DR-submodular optimization problems. The paper extends these results to multiple feedback settings, facilitating conversions between semi-bandit/first-order feedback and bandit/zeroth-order feedback, as well as between first/zeroth-order feedback and semi-bandit/bandit feedback. Leveraging this framework, new algorithms are derived using existing results as base algorithms for convex optimization, improving upon state-of-the-art results in various cases. Dynamic and adaptive regret guarantees are obtained for DR-submodular maximization, marking the first algorithms to achieve such guarantees in these settings. Notably, the paper achieves these advancements with fewer assumptions compared to existing state-of-the-art results, underscoring its broad applicability and theoretical contributions to non-convex optimization.

math.OC

Complexity of Integer Programming in Reverse Convex Sets via Boundary Hyperplane Cover

We study the complexity of identifying the integer feasibility of reverse convex sets. We present various settings where the complexity can be either NP-Hard or efficiently solvable when the dimension is fixed. Of particular interest is the case of bounded reverse convex constraints with a polyhedral domain. We introduce a structure, \emph{Boundary Hyperplane Cover}, that permits this problem to be solved in polynomial time in fixed dimension provided the number of nonlinear reverse convex sets is fixed.

math.OC

Size-Selective Threshold Harvesting under Nonlocal Crowding and Exogenous Recruitment

We propose a nonlinear size-structured fishery model for externally recruited stocks under size-selective harvesting. The population density satisfies a McKendrick--von Foerster transport equation in which growth and natural mortality depend on a nonlocal crowding index, while harvesting acts as a bounded size-dependent mortality control. Unlike standard self-recruiting formulations, recruitment is prescribed as a lower-boundary inflow, making the model suitable for enhancement fisheries or analyses conditional on juvenile input. For the no-harvest baseline, we derive the stationary size profile and reduce the nonlinear equilibrium problem to a scalar closure equation, proving existence and uniqueness under a net monotonicity condition. We introduce an intrinsic replacement index and show why, in this externally forced setting, it is a viability diagnostic rather than a persistence threshold. A formal state--adjoint system yields a bang--bang switching rule; under weak coupling and single crossing, the optimal policy has a threshold structure. Numerical experiments validate the approximation and sensitivity trends

math.OC