arXiv · 2001.09032
Limits on Gradient Compression for Stochastic Optimization
Abstract
We consider stochastic optimization over $\ell_p$ spaces using access to a first-order oracle. We ask: {What is the minimum precision required for oracle outputs to retain the unrestricted convergence rates?} We characterize this precision for every $p\geq 1$ by deriving information theoretic lower bounds and by providing quantizers that (almost) achieve these lower bounds. Our quantizers are new and easy to implement. In particular, our results are exact for $p=2$ and $p=\infty$, showing the minimum precision needed in these settings are $\Theta(d)$ and $\Theta(\log d)$, respectively. The latter result is surprising since recovering the gradient vector will require $\Omega(d)$ bits.
Explore related subjects
Keep this discovery
Prathamesh Mayekar, Himanshu Tyagi. 2020-01-24. Limits on Gradient Compression for Stochastic Optimization. https://arxiv.org/abs/2001.09032
Cite the original work for its findings. Save a collection to share your selection of sources.