Search arXiv⌕ Search

arXiv · 2009.03971

Fast adaptive by constants of strong-convexity and Lipschitz for gradient first order methods

Abstract

The work is devoted to the construction of efficient and applicable to real tasks first-order methods of convex optimization, that is, using only values of the target function and its derivatives. Construction uses OGM-G, fast gradient method which is optimal by complexity, but requires to know the Lipschitz constant for gradient and the strong convexity constant to determine the number of steps and step length. This requirement makes practical usage impossible. An adaptive on the constant for strong convexity algorithm ACGM is proposed, based on restarts of the OGM-G with update of the strong convexity constant estimate, and an adaptive on the Lipschitz constant for gradient ALGM, in which the use of OGM-G restarts is supplemented by the selection of the Lipschitz constant with verification of the convexity conditions used in the universal gradient descent method. This eliminates the disadvantages of the original method associated with the need to know these constants, which makes practical usage possible. Optimality of estimates for the complexity of the constructed algorithms is proved. To verify the results obtained, experiments on model functions and real tasks from machine learning are carried out.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nikita Pletnev. 2020-09-08. Fast adaptive by constants of strong-convexity and Lipschitz for gradient first order methods. https://arxiv.org/abs/2009.03971

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Discreteness to Convexity: Promotion Planning via Simplotope Triangulation

Price promotion optimization is a computationally challenging problem central to supermarket operations, requiring simultaneous pricing decisions across multiple products and periods. This paper introduces a new formulation for price promotion by developing convex hull results for supermodular compositions of univariate functions over a simplotope. Leveraging this reformulation with Gurobi, we achieve substantial performance gains: instances with up to 125 products, 20 periods, and 5 price levels are solved in an average of 7 minutes, demonstrating the potential to handle even larger instances. Our exact solution methods extract 25--48\% additional profit from promotion planning relative to state-of-the-art heuristic approaches. Additionally, we extend the polynomially solvable cases from two to multiple price levels and expand our results to allow for multiplicative historical effects. Our core methodological innovation applies to a broad class of nonlinear discrete optimization problems. Specifically, our results convexify a class of nonlinear functions that includes monomials and the widely studied L natural function structure.

math.OC↗

Deterministic Mean Field Games on Networks and Related Optimal Control Problems

We study a class of deterministic mean field games and related optimal control problems, with a finite time horizon and in which the state space is a network. An agent controls her velocity, and, when she occupies a vertex, she can either remain still or enter any adjacent edge. The running and terminal costs are assumed to be continuous in each edge, but may jump at the vertices. Compared to the companion paper [4], we make more general assumptions about the costs and consider networks with an arbitrary number of vertices; this higher degree of generality brings new difficulties. For the optimal control problems mentioned above, we obtain in particular the existence of optimal trajectories and regularity results concerning the optimal trajectories and the value function. These control theoretic results make it possible to address a class of mean field games on networks, with costs that do not depend separately on the control and on the distribution of states, and that are non-local with respect to the latter. Focusing on a Lagrangian formulation, we obtain the existence of relaxed equilibria consisting of probability measures on admissible trajectories. To any relaxed equilibrium corresponds a mild solution, i.e. a pair $(u, m)$ made of the value function $u$ of a related optimal control problem and a family $m = (m(t))_t$ of probability measures on the network. Given $m$, the value function $u$ is a viscosity solution of a Hamilton-Jacobi problem on the network. We then investigate the regularity properties of $u$ and a weak form of a Fokker-Planck equation satisfied by $m$.

math.OC↗

Disjunctive Submodular Functions: Envelopes and Applications to Inventory and 0-1 Quadratic Optimization

This paper considers convex envelopes of disjunctive submodular functions---functions that are lattice family submodular over faces of a hypercube---and constructs the first strongly polynomial algorithm for their separation when there are two facial disjunctions. Submodular functions, whose convex envelopes are characterized by the Lovász extension, have occupied a fundamental role in constructing relaxations for combinatorial and nonlinear optimization problems. However, disjunctive submodular function envelopes have not been explored besides the use of ellipsoid algorithm, which remains practically intractable. Our algorithm is derived in three steps by expressing the disjunctive function as a minimum of two extended submodular functions, introducing a variable lifting technique, and constructing the sublinear envelope in the lifted space. The paper also makes several other contributions. First, we provide a disjunctive formulation for the case where each submodular function admits a linear programming formulation. Second, we derive the closed-form sublinear envelope characterization for intersecting submodular functions, yielding new structural insights into a multi-product inventory sales maximization problem. Third, we fully characterize the convex envelope of a bilinear function defined over a cycle graph in the original variable space. Finally, we show computationally that the cycle inequalities close approximately 60\% of the gap for complete and Hadamard graphs, over 30\% of the gap for complete bipartite graphs, and over 80\% of the gap for sparse graphs such as cactus and Halin graphs. The resulting relaxations are also more efficient to solve than previous extended space formulations.

math.OC↗