2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) · 2022 · 18 citations · 29 references
Face Part DiscoverySubpart CapsulesEngineeringMachine LearningBiometricsCapsule NetworksFace DetectionImage ClassificationFacial Recognition SystemImage AnalysisData SciencePattern RecognitionVideo TransformerVision RecognitionMachine VisionFeature LearningVision Language ModelComputer ScienceMedical Image ComputingDeep LearningComputer VisionPart-level Capsules
Capsule networks are designed to present the objects by a set of parts and their relationships, which provide an insight into the procedure of visual perception. Although recent works have shown the success of capsule networks on simple objects like digits, the human faces with homologous structures, which are suitable for capsules to describe, have not been explored. In this paper, we propose a Hierarchical Parsing Capsule Network (HP-Capsule) for unsupervised face subpart-part discovery. When browsing large-scale face images without labels, the network first encodes the frequently observed patterns with a set of explainable subpart capsules. Then, the subpart capsules are assembled into part-level capsules through a Transformer-based Parsing Module (TPM) to learn the compositional relations between them. During training as the face hierarchy is progressively built and refined, the part capsules adaptively encode the face parts with semantic consistency. HP-Capsule extends the application of capsule networks from digits to human faces and takes a step forward to show how the neural networks understand homologous objects without human intervention. Besides, HP-Capsule gives unsupervised face segmentation results by the covered regions of part capsules, enabling qualitative and quantitative evaluation. Experiments on BP4D and Multi-PIE datasets show the effectiveness of our method.
29
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
Scikit-learn: Machine Learning in Python
Fabián Pedregosa, Gaël Varoquaux, Alexandre Gramfort et al. · arXiv (Cornell University) · 2012 · 63.3K citations · Full text
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu et al. · 2020 · 11.6K citations
Convolutional Neural Network, Image Analysis, Machine Learning +14