2009 · 25 citations · 30 references
Quality ExpansionsHighest Quality ExpansionsEngineeringStructured DataSemi-automatic EntitySemantic WebSemanticsText MiningNatural Language ProcessingInformation RetrievalData ScienceData MiningComputational LinguisticsManagementData IntegrationSemi-structured DataNamed-entity RecognitionData ManagementVery Large DatabaseEntity DisambiguationKnowledge DiscoveryComputer ScienceDatabase TheoryRefinement ModelsAutomated ReasoningFormal Methods
State of the art set expansion algorithms produce varying quality expansions for different entity types. Even for the highest quality expansions, errors still occur and manual refinements are necessary for most practical uses. In this paper, we propose algorithms to aide this refinement process, greatly reducing the amount of manual labor required. The methods rely on the fact that most expansion errors are systematic, often stemming from the fact that some seed elements are ambiguous. Using our methods, empirical evidence shows that average R-precision over random entity sets improves by 26% to 51% when given from 5 to 10 manually tagged errors. Both proposed refinement models have linear time complexity in set size allowing for practical online use in set expansion systems.
30
Automatic acquisition of hyponyms from large text corpora
Marti A. Hearst · 1992 · 3.3K citations · Full text
Automatic retrieval and clustering of similar words
Dekang Lin · 1998 · 1.6K citations · Full text
Open information extraction from the web
Michele Banko, Michael Cafarella, Stephen Soderland et al. · 2007 · 1.3K citations