Search arXiv⌕ Search

arXiv · 1807.06629

Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning

Abstract

In distributed training of deep neural networks, parallel mini-batch SGD is widely used to speed up the training process by using multiple workers. It uses multiple workers to sample local stochastic gradient in parallel, aggregates all gradients in a single server to obtain the average, and update each worker's local model using a SGD update with the averaged gradient. Ideally, parallel mini-batch SGD can achieve a linear speed-up of the training time (with respect to the number of workers) compared with SGD over a single worker. However, such linear scalability in practice is significantly limited by the growing demand for gradient communication as more workers are involved. Model averaging, which periodically averages individual models trained over parallel workers, is another common practice used for distributed training of deep neural networks since (Zinkevich et al. 2010) (McDonald, Hall, and Mann 2010). Compared with parallel mini-batch SGD, the communication overhead of model averaging is significantly reduced. Impressively, tremendous experimental works have verified that model averaging can still achieve a good speed-up of the training time as long as the averaging interval is carefully controlled. However, it remains a mystery in theory why such a simple heuristic works so well. This paper provides a thorough and rigorous theoretical study on why model averaging can work as well as parallel mini-batch SGD with significantly less communication overhead.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hao Yu, Sen Yang, Shenghuo Zhu. 2018-11-16. Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning. https://arxiv.org/abs/1807.06629

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Discreteness to Convexity: Promotion Planning via Simplotope Triangulation

Price promotion optimization is a computationally challenging problem central to supermarket operations, requiring simultaneous pricing decisions across multiple products and periods. This paper introduces a new formulation for price promotion by developing convex hull results for supermodular compositions of univariate functions over a simplotope. Leveraging this reformulation with Gurobi, we achieve substantial performance gains: instances with up to 125 products, 20 periods, and 5 price levels are solved in an average of 7 minutes, demonstrating the potential to handle even larger instances. Our exact solution methods extract 25--48\% additional profit from promotion planning relative to state-of-the-art heuristic approaches. Additionally, we extend the polynomially solvable cases from two to multiple price levels and expand our results to allow for multiplicative historical effects. Our core methodological innovation applies to a broad class of nonlinear discrete optimization problems. Specifically, our results convexify a class of nonlinear functions that includes monomials and the widely studied L natural function structure.

math.OC↗

Deterministic Mean Field Games on Networks and Related Optimal Control Problems

We study a class of deterministic mean field games and related optimal control problems, with a finite time horizon and in which the state space is a network. An agent controls her velocity, and, when she occupies a vertex, she can either remain still or enter any adjacent edge. The running and terminal costs are assumed to be continuous in each edge, but may jump at the vertices. Compared to the companion paper [4], we make more general assumptions about the costs and consider networks with an arbitrary number of vertices; this higher degree of generality brings new difficulties. For the optimal control problems mentioned above, we obtain in particular the existence of optimal trajectories and regularity results concerning the optimal trajectories and the value function. These control theoretic results make it possible to address a class of mean field games on networks, with costs that do not depend separately on the control and on the distribution of states, and that are non-local with respect to the latter. Focusing on a Lagrangian formulation, we obtain the existence of relaxed equilibria consisting of probability measures on admissible trajectories. To any relaxed equilibrium corresponds a mild solution, i.e. a pair $(u, m)$ made of the value function $u$ of a related optimal control problem and a family $m = (m(t))_t$ of probability measures on the network. Given $m$, the value function $u$ is a viscosity solution of a Hamilton-Jacobi problem on the network. We then investigate the regularity properties of $u$ and a weak form of a Fokker-Planck equation satisfied by $m$.

math.OC↗

Disjunctive Submodular Functions: Envelopes and Applications to Inventory and 0-1 Quadratic Optimization

This paper considers convex envelopes of disjunctive submodular functions---functions that are lattice family submodular over faces of a hypercube---and constructs the first strongly polynomial algorithm for their separation when there are two facial disjunctions. Submodular functions, whose convex envelopes are characterized by the Lovász extension, have occupied a fundamental role in constructing relaxations for combinatorial and nonlinear optimization problems. However, disjunctive submodular function envelopes have not been explored besides the use of ellipsoid algorithm, which remains practically intractable. Our algorithm is derived in three steps by expressing the disjunctive function as a minimum of two extended submodular functions, introducing a variable lifting technique, and constructing the sublinear envelope in the lifted space. The paper also makes several other contributions. First, we provide a disjunctive formulation for the case where each submodular function admits a linear programming formulation. Second, we derive the closed-form sublinear envelope characterization for intersecting submodular functions, yielding new structural insights into a multi-product inventory sales maximization problem. Third, we fully characterize the convex envelope of a bilinear function defined over a cycle graph in the original variable space. Finally, we show computationally that the cycle inequalities close approximately 60\% of the gap for complete and Hadamard graphs, over 30\% of the gap for complete bipartite graphs, and over 80\% of the gap for sparse graphs such as cactus and Halin graphs. The resulting relaxations are also more efficient to solve than previous extended space formulations.

math.OC↗