2021 · 34 citations · 49 references
EngineeringMachine LearningMultimodal LearningVideo SummarizationLanguage LearningVideo InterpretationLanguage UnderstandingSpeech RecognitionNatural Language ProcessingMultimodal LlmComputational LinguisticsLanguage StudiesPre-trained Neural ModelsMachine TranslationLinguisticsVision Language ModelMultimodal ContentVideo UnderstandingMultimodal TranslationDeep LearningComputer VisionChinese Video
The pre-trained neural models have recently achieved impressive performance in understanding multimodal content. However, it is still very challenging to pre-train neural models for video and language understanding, especially for Chinese video-language data, due to the following reasons. Firstly, existing video-language pre-training algorithms mainly focus on the co-occurrence of words and video frames, but ignore other valuable semantic and structure information of video-language content, e.g., sequential order and spatiotemporal relationships. Secondly, there exist conflicts between video sentence alignment and other proxy tasks. Thirdly, there is a lack of large-scale and high-quality Chinese video-language datasets (eg. including 10 million unique videos), which are the fundamental success conditions for pre-training techniques. In this work, we propose a novel video-language understanding framework named Victor, which stands for VIdeo-language understanding via Contrastive mulTimOdal pRe-training. Besides general proxy tasks such as masked language modeling, Victor constructs several novel proxy tasks under the contrastive learning paradigm, making the model be more robust and able to capture more complex multimodal semantic and structural relationships from different perspectives. Victor is trained on a large-scale Chinese video-language dataset, including over 10 million complete videos with corresponding high-quality textual descriptions. We apply the pre-trained Victor model to a series of downstream applications and demonstrate its superior performance, comparing against the state-of-the-art pre-training methods such as VideoBERT and UniVL.
49
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu et al. · 2020 · 11.6K citations
Convolutional Neural Network, Image Analysis, Machine Learning +14