IEEE Transactions on Circuits and Systems for Video Technology · 2021 · 41 citations · 47 references
Natural Language ProcessingMultimodal LlmStacked ArchitectureExposure BiasMultimodal Attention NetworkMachine LearningEngineeringVisual GroundingVision Language ModelVideo SummarizationPre-trained ModelsMultimodal ProcessingDeep LearningRecent Neural ModelsLanguage ProcessingComputer VisionMachine TranslationMulti-modal Summarization
Recent neural models for video captioning usually employ an attention-based encoder-decoder framework. However, current approaches mainly attend to the motion features and object features of the video when generating the caption, but ignore the potential but useful historical information. Besides, exposure bias and vanishing gradients problems always exist in current caption generation models. In this paper, we propose a novel video captioning framework, named Stacked Multimodal Attention Network (SMAN). It adopts additional visual and textual historical information during caption generation as context features, employs a stacked architecture to process different features gradually, and utilizes the Reinforcement Learning method and coarse-to-fine training strategy to further improve the generated results. Both quantitative and qualitative experiments on the benchmark datasets of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">MSVD</i> and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">MSR-VTT</i> show the effectiveness and feasibility of our framework. The codes are available on <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/zhengyi123456/SMAN</uri> .
47
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7
Kishore Papineni, Salim Roukos, Todd J. Ward et al. · 2001 · 20.9K citations · Full text
Natural Language Processing, Computer-assisted Translation, Engineering +10
Aggregated Residual Transformations for Deep Neural Networks
Saining Xie, Ross Girshick, Piotr Dollár et al. · 2017 · 11.6K citations
Convolutional Neural Network, Engineering, Machine Learning +16