arXiv (Cornell University) · 2021 · 10 citations · 0 references
Pruning neural networks reduces inference time and memory costs. On standard\nhardware, these benefits will be especially prominent if coarse-grained\nstructures, like feature maps, are pruned. We devise two novel saliency-based\nmethods for second-order structured pruning (SOSP) which include correlations\namong all structures and layers. Our main method SOSP-H employs an innovative\nsecond-order approximation, which enables saliency evaluations by fast\nHessian-vector products. SOSP-H thereby scales like a first-order method\ndespite taking into account the full Hessian. We validate SOSP-H by comparing\nit to our second method SOSP-I that uses a well-established Hessian\napproximation, and to numerous state-of-the-art methods. While SOSP-H performs\non par or better in terms of accuracy, it has clear advantages in terms of\nscalability and efficiency. This allowed us to scale SOSP-H to large-scale\nvision tasks, even though it captures correlations across all layers of the\nnetwork. To underscore the global nature of our pruning methods, we evaluate\ntheir performance not only by removing structures from a pretrained network,\nbut also by detecting architectural bottlenecks. We show that our algorithms\nallow to systematically reveal architectural bottlenecks, which we then remove\nto further increase the accuracy of the networks.\n