IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022 · 127 citations · 41 references
Convolutional Neural NetworkParameter SpaceEngineeringMachine LearningImage ClassificationImage AnalysisData SciencePattern RecognitionResidual BranchesResidual LearningLong-tail LearningVideo TransformerSemi-supervised LearningData AugmentationMachine VisionFeature LearningTail ClassesComputer ScienceDeep LearningComputer VisionLimited Data Learning
Deep learning struggles with long‑tailed data distributions, a common real‑world issue. The study proposes a parameter‑space approach to preserve capacity for low‑frequency classes. They introduce a residual fusion network with a main branch for all classes and two residual branches that progressively enhance medium‑ and tail‑class performance, then aggregate via additive shortcuts. Experiments on long‑tailed CIFAR, Places, ImageNet, and iNaturalist show the method’s effectiveness. Code is available at https://github.com/jiequancui/ResLT.
Deep learning algorithms face great challenges with long-tailed data distribution which, however, is quite a common case in real-world scenarios. Previous methods tackle the problem from either the aspect of input space (re-sampling classes with different frequencies) or loss space (re-weighting classes with different weights), suffering from heavy over-fitting to tail classes or hard optimization during training. To alleviate these issues, we propose a more fundamental perspective for long-tailed recognition, i.e., from the aspect of parameter space, and aims to preserve specific capacity for classes with low frequencies. From this perspective, the trivial solution utilizes different branches for the head, medium, tail classes respectively, and then sums their outputs as the final results is not feasible. Instead, we design the effective residual fusion mechanism - with one main branch optimized to recognize images from all classes, another two residual branches are gradually fused and optimized to enhance images from medium+tail classes and tail classes respectively. Then the branches are aggregated into final results by additive shortcuts. We test our method on several benchmarks, i.e., long-tailed version of CIFAR-10, CIFAR-100, Places, ImageNet, and iNaturalist 2018. Experimental results manifest the effectiveness of our method. Our code is available at https://github.com/jiequancui/ResLT.
41
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia et al. · 2015 · 46.2K citations
Image Classification, Deep Neural Networks, Image Analysis +15
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7
SMOTE: Synthetic Minority Over-sampling Technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence Hall et al. · Journal of Artificial Intelligence Research · 2002 · 29.6K citations · Full text