arXiv (Cornell University) · 2018 · 74 citations · 10 references
EngineeringEntity SummarizationCorpus LinguisticsText MiningAutomatic SummarizationNatural Language ProcessingInformation RetrievalData ScienceComputational LinguisticsMulti- Document SummarizationLanguage StudiesContent AnalysisMachine TranslationSequence ModellingExtractive SummarizationKnowledge DiscoveryLong SequencesMulti-modal SummarizationRetrieval Augmented GenerationEnglish Wikipedia ArticlesLinguisticsLanguage Generation
We show that generating English Wikipedia articles can be approached as a multi- document summarization of source documents. We use extractive summarization to coarsely identify salient information and a neural abstractive model to generate the article. For the abstractive model, we introduce a decoder-only architecture that can scalably attend to very long sequences, much longer than typical encoder- decoder architectures used in sequence transduction. We show that this model can generate fluent, coherent multi-sentence paragraphs and even whole Wikipedia articles. When given reference documents, we show it can extract relevant factual information as reflected in perplexity, ROUGE scores and human evaluations.
10
The PageRank Citation Ranking : Bringing Order to the Web
Lawrence M. Page, Sergey Brin, Rajeev Motwani et al. · 1999 · 12.6K citations
Ashish Vaswani, Noam Shazeer, Niki Parmar et al. · 2025 · 6.5K citations · Full text
DBpedia – A large-scale, multilingual knowledge base extracted from Wikipedia
Jens Lehmann, Robert Isele, Max Jakob et al. · Semantic Web · 2015 · 3.2K citations · Full text