IEEE Transactions on Circuits and Systems for Video Technology · 2021 · 44 citations · 50 references
Few-shot LearningEngineeringMachine LearningGlobal InformationNatural Language ProcessingMultimodal LlmImage AnalysisZero-shot LearningData SciencePattern RecognitionGlobal FeaturesMachine TranslationMachine VisionFeature LearningSemantic AlignmentGlobal-local InterplayComputer ScienceComputer VisionLinguistics
Few-shot learning aims to recognize novel classes from only a few labeled training examples. Aligning semantically relevant local regions has shown promise in effectively comparing a query image with support images. However, global information is usually overlooked in the existing approaches, resulting in a higher possibility of learning semantics unrelated to the global information. To address this issue, we propose a Global-Local Interplay Metric Learning (GLIML) framework to employ the interplay between global features and local features to guide semantic alignment. We first design a Global-Local Information Concurrent Learning (GLICL) module to extract both global features and local features and perform global-local interplay. We then design a Global-Local Information Cross-Covariance Estimator (GLICCE) to learn the similarity on the global-local interplay, in contrast to the current practice where only local features are considered. Visualizations show that the global-local interplay decreases (1) the weights placed on the semantics that are irrelevant to the global information and (2) the variability of the learned features within every class in the feature space. Quantitative experiments on three benchmark datasets demonstrate that GLIML achieves state-of-the-art performance while maintaining high efficiency.
50
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Kaiming He, Georgia Gkioxari, Piotr Dollár et al. · 2017 · 27.9K citations
Object Instance Segmentation, Scene Analysis, Machine Vision +13
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu et al. · 2020 · 11.6K citations
Convolutional Neural Network, Image Analysis, Machine Learning +14
Xiaolong Wang, Ross Girshick, Abhinav Gupta et al. · 2018 · 11K citations
Prototypical Networks for Few-shot Learning
Jake Snell, Kevin Swersky, Richard S. Zemel · arXiv (Cornell University) · 2017 · 5.2K citations · Full text