Lancaster EPrints (Lancaster University) · 2006 · 28 citations · 16 references
Open access
In this paper, we report on our experi-ment to extract Chinese multiword ex-pressions from corpus resources as part of a larger research effort to improve a machine translation (MT) system. For ex-isting MT systems, the issue of multi-word expression (MWE) identification and accurate interpretation from source to target language remains an unsolved problem. Our initial test on the Chinese-to-English translation functions of Systran and CCID’s Huan-Yu-Tong MT systems reveal that, where MWEs are in-volved, MT tools suffer in terms of both comprehensibility and adequacy of the translated texts. For MT systems to be-come of further practical use, they need to be enhanced with MWE processing capability. As part of our study towards this goal, we test and evaluate a statistical tool, which was developed for English, for identifying and extracting Chinese MWEs. In our evaluation, the tool achieved precisions ranging from 61.16% to 93.96 % for different types of MWEs. Such results demonstrate that it is feasi-ble to automatically identify many Chi-nese MWEs using our tool, although it needs further improvement. 1
16
Accurate methods for the statistics of surprise and coincidence
Ted Dunning · 1993 · 2.7K citations
Engineering, Statistical Foundation, Rare Event Estimation +22
Stochastic inversion transduction grammars and bilingual parsing of parallel corpora
Dekai Wu · 1997 · 861 citations
Retrieving collocations from text: Xtract
Frank Smadja · 1993 · 799 citations
An empirical model of multiword expression decomposability
Timothy Baldwin, Colin Bannard, Takaaki Tanaka et al. · 2003 · 233 citations · Full text
Multiword Expression, Engineering, Part-of-speech Tagging +18
Ido Dagan, Kenneth Church · 1994 · 223 citations · Full text