arXiv (Cornell University) · 2020 · 22 citations · 18 references
How to generate descriptions from structured data organized in tables? Existing approaches using neural encoder-decoder models often suffer from lacking diversity. We claim that an open set of templates is crucial for enriching the phrase constructions and realizing varied generations. Learning such templates is prohibitive since it often requires a large paired corpus, which is seldom available. This paper explores the problem of automatically learning reusable "templates" from paired and non-paired data. We propose the variational template machine (VTM), a novel method to generate text descriptions from data tables. Our contributions include: a) we carefully devise a specific model architecture and losses to explicitly disentangle text template and semantic content information, in the latent spaces, and b)we utilize both small parallel data and large raw text without aligned tables to enrich the template learning. Experiments on datasets from a variety of different domains show that VTM is able to generate more diversely while keeping a good fluency and quality.
18
The Curious Case of Neural Text Degeneration
Ari Holtzman, Jan Buys, Leo Du et al. · arXiv (Cornell University) · 2020 · 527 citations · Full text
Challenges in Data-to-Document Generation
Sam Wiseman, Stuart M. Shieber, Alexander M. Rush · 2017 · 491 citations · Full text
Style Transfer from Non-Parallel Text by Cross-Alignment
Tianxiao Shen, Tao Leí, Regina Barzilay et al. · arXiv (Cornell University) · 2017 · 442 citations · Full text
Natural Language Processing, Engineering, Sentiment Modification +11
Yaoming Zhu, Sidi Lu, Lei Zheng et al. · 2018 · 387 citations
Data Generation, Natural Language Processing, Retrieval Augmented Generation +14