Proceedings of the AAAI Conference on Artificial Intelligence · 2021 · 33 citations · 62 references
Artificial IntelligenceEngineeringMachine LearningShortcut EffectsCognitionSemanticsSocial SciencesLanguage PriorsNatural Language ProcessingVisual CommonsenseMultimodal LlmVisual GroundingData ScienceCommonsense KnowledgeVisual Question AnsweringVisual Commonsense ReasoningCognitive ScienceReasoning SystemSemantic InterpretationCommonsense ReasoningVision Language ModelComputer ScienceVisual ReasoningAutomated ReasoningCase StudyHuman-computer Interaction
Visual reasoning and question-answering have gathered attention in recent years. Many datasets and evaluation protocols have been proposed; some have been shown to contain bias that allows models to ``cheat'' without performing true, generalizable reasoning. A well-known bias is dependence on language priors (frequency of answers) resulting in the model not looking at the image. We discover a new type of bias in the Visual Commonsense Reasoning (VCR) dataset. In particular we show that most state-of-the-art models exploit co-occurring text between input (question) and output (answer options), and rely on only a few pieces of information in the candidate options, to make a decision. Unfortunately, relying on such superficial evidence causes models to be very fragile. To measure fragility, we propose two ways to modify the validation data, in which a few words in the answer choices are modified without significant changes in meaning. We find such insignificant changes cause models' performance to degrade significantly. To resolve the issue, we propose a curriculum-based masking approach, as a mechanism to perform more robust training. Our method improves the baseline by requiring it to pay attention to the answers as a whole, and is more effective than prior masking strategies.
62
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky et al. · 2014 · 34.2K citations
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, Christopher D. Manning · 2014 · 33.2K citations