arXiv (Cornell University) · 2019 · 88 citations · 33 references
Convolutional Neural NetworkEngineeringMachine LearningWeight DecayRecurrent Neural NetworkSpeech RecognitionNatural Language ProcessingData ScienceSparse Neural NetworkSupervised LearningMachine TranslationLayer-wise Gradient NormalizationMachine Learning ModelStochastic Gradient MethodsComputer ScienceDeep LearningNeural Architecture SearchSpeech ProcessingLayer-wise Adaptive MomentsDeep Networks
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with momentum and Adam or AdamW. Additionally, NovoGrad (1) is robust to the choice of learning rate and weight initialization, (2) works well in a large batch setting, and (3) has two times smaller memory footprint than Adam.
33
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe et al. · 2016 · 30.2K citations
Convolutional Neural Network, Engineering, Machine Learning +17
Decoupled Weight Decay Regularization
Ilya Loshchilov, Frank Hutter · arXiv (Cornell University) · 2017 · 9K citations · Full text
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John C. Duchi, Elad Hazan, Yoram Singer · 2010 · 8.6K citations
Ashish Vaswani, Noam Shazeer, Niki Parmar et al. · 2025 · 6.5K citations · Full text