2000 · 72 citations · 13 references
One of the approaches to cross-language information retrieval (CLIR) is based on the use of parallel texts. In this paper, we will describe a parallel text mining system called PTMiner (Parallel Text Miner) for the Web environment. We will explain the underlying mining algorithm of this system as well as its implementation using a distributed model and database technology. The resulted corpora are used as the training material for statistical translation models. Preliminary experimental results using the models for CLIR are reported. 1 Introduction Data mining, text mining and other knowledge discovering techniques have become an attractive research area in the past years. The enormous amount of information often oers potential solutions to some problems. This is the case of parallel texts that provide translation examples from a language to another. A pair of parallel texts is two such texts that are translation one for the other. In our work, the need for parallel corpora is orig...
13
The mathematics of statistical machine translation: parameter estimation
Peter F. Brown, Vincent J. Della Pietra, Stephen A. Della Pietra et al. · 1993 · 4.1K citations
Aligning sentences in parallel corpora
Peter F. Brown, Jennifer C. Lai, Robert L. Mercer · 1991 · 487 citations · Full text
The Harvest information discovery and access system
C. Mic Bowman, Peter B. Danzig, Darren Hardy et al. · Computer Networks and ISDN Systems · 1995 · 421 citations