arXiv (Cornell University) · 2020 · 17 citations · 15 references
Artificial IntelligenceEngineeringMachine LearningNetwork AnalysisData ScienceData MiningSparse Neural NetworkCombinatorial OptimizationScaling LawNeural Scaling LawScaling AnalysisMagnitude-pruned NetworksMachine Learning ModelPruning Across ScalesKnowledge DiscoveryComputer ScienceDeep LearningNeural Architecture SearchPruned NetworksFeature ScalingModel Compression
We show that the error of iteratively magnitude-pruned networks empirically follows a scaling law with interpretable coefficients that depend on the architecture and task. We functionally approximate the error of the pruned networks, showing it is predictable in terms of an invariant tying width, depth, and pruning level, such that networks of vastly different pruned densities are interchangeable. We demonstrate the accuracy of this approximation over orders of magnitude in depth, width, dataset size, and density. We show that the functional form holds (generalizes) for large scale data (e.g., ImageNet) and architectures (e.g., ResNets). As neural networks become ever larger and costlier to train, our findings suggest a framework for reasoning conceptually and analytically about a standard method for unstructured pruning.
15
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Language Models are Few-Shot Learners
T. B. Brown, Benjamin Mann, Nick Ryder et al. · arXiv (Cornell University) · 2020 · 3K citations · Full text
Artificial Intelligence, Few-shot Learning, Llm Fine-tuning +20
Yann LeCun, John S. Denker, Sara A. Solla · 1989 · 2.6K citations
Russell Reed · IEEE Transactions on Neural Networks · 1993 · 1.7K citations
Artificial Intelligence, Incremental Learning, Engineering +21