arXiv · 2009.11224
Applying the Roofline model for Deep Learning performance optimizations
Abstract
In this paper We present a methodology for creating Roofline models automatically for Non-Unified Memory Access (NUMA) using Intel Xeon as an example. Finally, we present an evaluation of highly efficient deep learning primitives as implemented in the Intel oneDNN Library.
Explore related subjects
Keep this discovery
Jacek Czaja, Michal Gallus, Joanna Wozna, Adam Grygielski, Luo Tao. 2020-09-23. Applying the Roofline model for Deep Learning performance optimizations. https://arxiv.org/abs/2009.11224
Cite the original work for its findings. Save a collection to share your selection of sources.