1998 · 149 citations · 24 references
Ontology (Information Science)Structured VocabularyEngineeringKnowledge ExtractionSemantic WebSemanticsCorpus LinguisticsText MiningNatural Language ProcessingInformation RetrievalData ScienceData IntegrationOntology LearningApplication OntologyKnowledge DiscoveryTerminology ExtractionUnstructured DocumentInformation ExtractionUnstructured DocumentsOntology-based ExtractionData Extraction
We present a new approach to extracting information from unstructured documents based on an application ontology that describes a domain of interest. Starting with such an ontology, we formulate rules to extract constants and context keywords from unstructured documents. For each unstructured document of interest, we extract its constants and keywords and apply a recognizer to organize extracted constants as attribute values of tuples in a generated database schema. To make our approach general, we fix all the processes and change only the ontological description for a different application domain. In experiments we conducted on two different types of unstructured documents taken from the Web, our approach attained recall ratios in the 80% and 90% range and precision ratios near 98%.
24
The Lorel query language for semistructured data
Serge Abiteboul, Dallan Quass, Jason McHugh et al. · International Journal on Digital Libraries · 1997 · 1.1K citations
Wrapper induction for information extraction
Nicholas Kushmerick, Daniel S. Weld · 1997 · 1K citations
Jim Cowie, Wendy G. Lehnert · Communications of the ACM · 1996 · 729 citations · Full text