2012 · 16 citations · 9 references
EngineeringSemanticsSemantic WebCorpus LinguisticsText MiningNatural Language ProcessingInformation RetrievalComputational LinguisticsText SimplificationLimited-size LexiconLanguage StudiesWikipedia ArticlesMachine TranslationAutomatic GenerationComputational LexicologyWordnet-based Lexical SimplificationTerminology ExtractionDistributional SemanticsLexical ResourceKeyword ExtractionLinguisticsWord-sense Disambiguation
We explore algorithms for the automatic generation of a limited-size lexicon from a document, such that the lexicon covers as much as possible of the semantic space of the original document, as specifically as possible. We evaluate six related algorithms that automatically derive limited-size vocabularies from Wikipedia articles, focusing on nouns and verbs. The proposed algorithms combine Personalized Page Rank (Agirre and Soroa, 2009) and principles of information maximization, beginning with a user-supplied document and constructing a customized small vocabulary using WordNet. The bestperforming algorithm relies on word-sense disambiguation with sentence-level context information at the earliest stage of analysis, indicating that this computationally costly task is nonetheless valuable.
9
Personalizing PageRank for word sense disambiguation
Eneko Agirre, Aitor Soroa · 2009 · 534 citations · Full text
Simple English Wikipedia: A New Text Simplification Task
W. Cöster, David Kauchak · 2011 · 181 citations