IEEE Access · 2020 · 28 citations · 53 references
EngineeringMachine LearningNatural Language ProcessingMultimodal LlmImage AnalysisText-to-image RetrievalData ScienceVisual GroundingComputational LinguisticsImage Caption GenerationFlickr8k Cn DatasetVisual Question AnsweringLanguage StudiesMachine TranslationVision Language ModelDeep LearningComputer VisionVisual Attention ModelImage CaptionLinguistics
As an interesting and challenging problem, generating image caption automatically has attracted increasingly attention in natural language processing and computer vision communities. In this paper, we propose an end-to-end deep learning approach for image caption generation. We leverage image feature information at specific location every moment and generate the corresponding caption description through a semantic attention model. The end-to-end framework allows us to introduce an independent recurrent structure as an attention module, derived by calculating the similarity between image feature sequence and semantic word sequence. Additionally, our model is designed to transfer the knowledge representation obtained from the English portion into the Chinese portion to achieve the cross-lingual image captioning. We evaluate the proposed model on the most popular benchmark datasets. We report an improvement of 3.9% over existing state-of-the-art approaches for cross-lingual image captioning on the Flickr8k CN dataset on CIDEr metric. The experimental results demonstrate the effectiveness of our attention model.
53
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, Quoc V. Le · arXiv (Cornell University) · 2014 · 13.3K citations · Full text
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin · 2004 · 8.3K citations
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros et al. · arXiv (Cornell University) · 2015 · 7.5K citations · Full text