Journal of the Physical Society of Japan · 2017 · 10 citations · 7 references
Artificial IntelligenceIntelligent Information ProcessingEngineeringMachine LearningNeural NetworkAlgorithmic LearningOnline LearningData ScienceWeight NormalizationRobot LearningLearning ProblemComputational Learning TheoryWeight VectorMachine Learning ModelLarge Scale OptimizationLearning AnalyticsComputer ScienceStatistical Learning TheoryDeep LearningNeural Architecture SearchModel OptimizationStatistical Mechanical Analysis
Weight normalization, a newly proposed optimization method for neural networks by Salimans and Kingma (2016), decomposes the weight vector of a neural network into a radial length and a direction vector, and the decomposed parameters follow their steepest descent update. They reported that learning with the weight normalization achieves better converging speed in several tasks including image recognition and reinforcement learning than learning with the conventional parameterization. However, it remains theoretically uncovered how the weight normalization improves the converging speed. In this study, we applied a statistical mechanical technique to analyze on-line learning in single layer linear and nonlinear perceptrons with weight normalization. By deriving order parameters of the learning dynamics, we confirmed quantitatively that weight normalization realizes fast converging speed by automatically tuning the effective learning rate, regardless of the nonlinearity of the neural network. This property is realized when the initial value of the radial length is near the global minimum; therefore, our theory suggests that it is important to choose the initial value of the radial length appropriately when using weight normalization.
7
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John C. Duchi, Elad Hazan, Yoram Singer · 2010 · 8.6K citations
Natural Gradient Works Efficiently in Learning
Neural Computation · 1998 · 2.7K citations