arXiv · 2006.07253
Dynamic Model Pruning with Feedback
Abstract
Deep neural networks often have millions of parameters. This can hinder their deployment to low-end devices, not only due to high memory requirements but also because of increased latency at inference. We propose a novel model compression method that generates a sparse trained model without additional overhead: by allowing (i) dynamic allocation of the sparsity pattern and (ii) incorporating feedback signal to reactivate prematurely pruned weights we obtain a performant sparse model in one single training pass (retraining is not needed, but can further improve the performance). We evaluate our method on CIFAR-10 and ImageNet, and show that the obtained sparse models can reach the state-of-the-art performance of dense models. Moreover, their performance surpasses that of models generated by all previously proposed pruning schemes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev, Martin Jaggi. 2020-06-12. Dynamic Model Pruning with Feedback. https://arxiv.org/abs/2006.07253
Cite the original work for its findings. Save a collection to share your selection of sources.