arXiv · 1904.11943
SWALP : Stochastic Weight Averaging in Low-Precision Training
Abstract
Low precision operations can provide scalability, memory savings, portability, and energy efficiency. This paper proposes SWALP, an approach to low precision training that averages low-precision SGD iterates with a modified learning rate schedule. SWALP is easy to implement and can match the performance of full-precision SGD even with all numbers quantized down to 8 bits, including the gradient accumulators. Additionally, we show that SWALP converges arbitrarily close to the optimal solution for quadratic objectives, and to a noise ball asymptotically smaller than low precision SGD in strongly convex settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, Christopher De Sa. 2019-05-20. SWALP : Stochastic Weight Averaging in Low-Precision Training. https://arxiv.org/abs/1904.11943
Cite the original work for its findings. Save a collection to share your selection of sources.