Journal of Computational and Graphical Statistics · 2023 · 14 citations · 13 references
AbstractWe study here a fixed mini-batch gradient descent (FMGD) algorithm to solve optimization problems with massive datasets. In FMGD, the whole sample is split into multiple nonoverlapping partitions. Once the partitions are formed, they are then fixed throughout the rest of the algorithm. For convenience, we refer to the fixed partitions as fixed mini-batches. Then for each computation iteration, the gradients are sequentially calculated on each fixed mini-batch. Because the size of fixed mini-batches is typically much smaller than the whole sample size, it can be easily computed. This leads to much reduced computation cost for each computational iteration. It makes FMGD computationally efficient and practically more feasible. To demonstrate the theoretical properties of FMGD, we start with a linear regression model with a constant learning rate. We study its numerical convergence and statistical efficiency properties. We find that sufficiently small learning rates are necessarily required for both numerical convergence and statistical efficiency. Nevertheless, an extremely small learning rate might lead to painfully slow numerical convergence. To solve the problem, a diminishing learning rate scheduling strategy can be used. This leads to the FMGD estimator with faster numerical convergence and better statistical efficiency. Finally, the FMGD algorithms with random shuffling and a general loss function are also studied. Supplementary materials for this article are available online.KEYWORDS: Fixed mini-batchGradient descentLearning rate schedulingRandom shufflingStochastic gradient descent Supplementary MaterialsThe detailed proofs of all propositions and theorems can be found in the supplementary materials.AcknowledgmentsThe authors thank the Editor, Associate Editor, and two anonymous reviewers for their careful reading and constructive comments.Disclosure StatementThe authors report there are no competing interests to declare.Additional informationFundingThis work is supported by National Natural Science Foundation of China (No.72001205), the Fundamental Research Funds for the Central Universities and the Research Funds of Renmin University of China (21XNA026).
13
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John C. Duchi, Elad Hazan, Yoram Singer · 2010 · 8.6K citations
Robust Stochastic Approximation Approach to Stochastic Programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan et al. · SIAM Journal on Optimization · 2009 · 2.1K citations