2018 · 170 citations · 32 references
Visual captioning has advanced, yet generating abstract stories from photo streams remains underexplored, challenged by expressive language, imaginary concepts, and limited evaluation metrics that impede reinforcement learning. The study aims to learn an implicit reward function from human demonstrations to enhance story generation. This is achieved by optimizing policy search within an Adversarial REward Learning (AREL) framework that learns the reward function. Automatic metrics show only a slight improvement over state‑of‑the‑art, but human evaluation reveals a significant increase in producing more human‑like stories.
Though impressive results have been achieved in visual captioning, the task of generating abstract stories from photo streams is still a little-tapped problem. Different from captions, stories have more expressive language styles and contain many imaginary concepts that do not appear in the images. Thus it poses challenges to behavioral cloning algorithms. Furthermore, due to the limitations of automatic metrics on evaluating story quality, reinforcement learning methods with hand-crafted rewards also face difficulties in gaining an overall performance boost. Therefore, we propose an Adversarial REward Learning (AREL) framework to learn an implicit reward function from human demonstrations, and then optimize policy search with the learned reward function. Though automatic evaluation indicates slight performance boost over state-of-the-art (SOTA) methods in cloning expert behaviors, human evaluation shows that our approach achieves significant improvement in generating more human-like stories than SOTA systems.
32
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Kishore Papineni, Salim Roukos, Todd J. Ward et al. · 2001 · 20.9K citations · Full text
Natural Language Processing, Computer-assisted Translation, Engineering +10
Convolutional Neural Networks for Sentence Classification
Yoon Kim · 2014 · 13.5K citations · Full text
Natural Language Processing, Llm Fine-tuning, Natural Language +14
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin · 2004 · 8.3K citations