Proceedings of the 30th ACM International Conference on Multimedia · 2022 · 23 citations · 29 references
Convolutional Neural NetworkEngineeringMachine LearningAffective NeurosciencePsychologySocial SciencesCausal InferenceEmotional ResponseData ScienceAffective ComputingData AugmentationCognitive ScienceFeature LearningVisual Emotion RecognitionCausal InterventionDeep LearningBackdoor AdjustmentFacial Expression RecognitionEmotionEmotion Recognition
Although much progress has been made in visual emotion recognition, researchers have realized that modern deep networks tend to exploit dataset characteristics to learn spurious statistical associations between the input and the target. Such dataset characteristics are usually treated as dataset bias, which damages the robustness and generalization performance of these recognition systems. In this work, we scrutinize this problem from the perspective of causal inference, where such dataset characteristic is termed as a confounder which misleads the system to learn the spurious correlation. To alleviate the negative effects brought by the dataset bias, we propose a novel Interventional Emotion Recognition Network (IERN) to achieve the backdoor adjustment, which is one fundamental deconfounding technique in causal inference. Specifically, IERN starts by disentangling the dataset-related context feature from the actual emotion feature, where the former forms the confounder. The emotion feature will then be forced to see each confounder stratum equally before being fed into the classifier. A series of designed tests validate the efficacy of IERN, and experiments on three emotion benchmarks demonstrate that IERN outperforms state-of-the-art approaches for unbiased visual emotion recognition.
29
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, Christopher D. Manning · 2014 · 33.2K citations
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa et al. · arXiv (Cornell University) · 2019 · 16.2K citations · Full text