2008 · 33 citations · 39 references
EngineeringSemantic WebSemanticsCorpus LinguisticsSemantic WikiText MiningNatural Language ProcessingInformation RetrievalData ScienceComputational LinguisticsOntology LearningSpecific OntologyLanguage StudiesNamed-entity RecognitionWikipedia DataEntity DisambiguationKnowledge DiscoveryTerminology ExtractionNamed EntitiesRelationship ExtractionLinguisticsWord-sense Disambiguation
Currently, for named entity disambiguation, the short-age of training data is a problem. This paper presents a novel method that overcomes this problem by automatically generating an annotated corpus based on a specific ontology. Then the corpus was enriched with new and informative features extracted from Wikipedia data. Moreover, rather than pursuing rule-based methods as in literature, we employ a machine learning model to not only disambiguate but also identify named entities. In addition, our method explores in details the use of a range of features extracted from texts, a given ontology, and Wikipedia data for disambiguation. This paper also systematically analyzes impacts of the features on disambiguation accuracy by varying their combinations for representing named entities. Empirical evaluation shows that, while the ontology provides basic features of named entities, Wikipedia is a fertile source for additional features to construct accurate and robust named entity disambiguation systems.
39
Zellig S. Harris · WORD · 1954 · 4.6K citations
Fabian M. Suchanek, Gjergji Kasneci, Gerhard Weikum · 2007 · 3.9K citations · Full text
Natural Language Processing, Extensible Ontology, Knowledge Base +14
Automatic word sense discrimination
Hinrich Schütze · 1998 · 1.3K citations