2009 · 23 citations · 24 references
EngineeringPart-of-speech TaggingCorpus LinguisticsText MiningNatural Language ProcessingApplied LinguisticsLanguage DocumentationData ScienceComputational LinguisticsGrammarAutomatic IdentificationLanguage StudiesNamed-entity RecognitionMachine TranslationNlp TaskKnowledge DiscoveryTerminology ExtractionMwe CandidatesMultiword ExpressionsTechnical DomainsLanguage CorpusLinguistics
Multiword Expressions (MWEs) are one of the stumbling blocks for more precise Natural Language Processing (NLP) systems. Particularly, the lack of coverage of MWEs in resources can impact negatively on the performance of tasks and applications, and can lead to loss of information or communication errors. This is especially problematic in technical domains, where a significant portion of the vocabulary is composed of MWEs. This paper investigates the use of a statistically-driven alignment-based approach to the identification of MWEs in technical corpora. We look at the use of several sources of data, including parallel corpora, using English and Portuguese data from a corpus of Pediatrics, and examining how a second language can provide relevant cues for this tasks. We report results obtained by a combination of statistical measures and linguistic information, and compare these to the reported in the literature. Such an approach to the (semi-)automatic identification of MWEs can considerably speed up lexicographic work, providing a more targeted list of MWE candidates.
24
Numerical recipes in C: the art of scientific computing
Choice Reviews Online · 1993 · 18K citations · Full text
Assessing agreement on classification tasks: the kappa statistic
Jean Carletta · ArXiv.org · 1996 · 2.1K citations · Full text
Roger K. Moore · 1986 · 637 citations