Across Languages and Cultures · 2019 · 34 citations · 14 references
Automatic Terminology ExtractionTerminology ManagementHeart FailureEngineeringLarge Crawled CorpusCorpus LinguisticsLanguage ProcessingText MiningNatural Language ProcessingInformation RetrievalData ScienceComputational LinguisticsTerm ExtractionCorpus AnalysisLanguage StudiesBiomedical Text MiningClinical LanguageLearner Corpus LinguisticsTerminology ExtractionMedical Terminology ExtractionMedical Language ProcessingLexical ResourceKeyword ExtractionMedicineLinguisticsHealth Informatics
Abstract We investigate the cost-effectiveness of special-purpose crawled corpora versus more focused corpora for automatic terminology extraction (ATE). Our focus is on medical terminology on heart failure for two languages, viz. English for which we have more web and specialized resources at our disposal and the less resourced Dutch. We show that, although term density in the dedicated corpora is larger for both languages, the potential for term extraction is higher in the crawled corpora than in the dedicated corpora. Furthermore, in a set of experiments in which we evaluate both types of corpora, while keeping size constant, we observe that more Gold Standard (GS) terms are covered by the “noisy” crawled corpus than with a dedicated corpus of the same size.
14
Introduction to the Special Issue on the Web as Corpus
Adam Kilgarriff, Gregory Grefenstette · Computational Linguistics · 2003 · 920 citations · Full text
Engineering, Semantic Web, Semantics +22
BootCaT: Bootstrapping corpora and terms from the web
Marco Baroni, Silvia Bernardini · 2004 · 376 citations