2006 · 66 citations · 14 references
Natural Language ProcessingComputer-assisted TranslationEngineeringIntermediate Phonemic MappingCorpus LinguisticsBengali WordsComputational LinguisticsLinguisticsLanguage TechnologyLanguage RecognitionArabic OrthographyNeural Machine TranslationSpeech ProcessingLanguage StudiesCollocational StatisticsSpeech TranslationMachine TranslationSpeech Recognition
Most machine transliteration systems transliterate out of vocabulary (OOV) words through intermediate phonemic mapping. A framework has been presented that allows direct orthographical mapping between two languages that are of different origins employing different alphabet sets. A modified joint source-channel model along with a number of alternatives have been proposed. Aligned transliteration units along with their context are automatically derived from a bilingual training corpus to generate the collocational statistics. The transliteration units in Bengali words take the pattern C+M where C represents a vowel or a consonant or a conjunct and M represents the vowel modifier or matra. The English transliteration units are of the form C*V* where C represents a consonant and V represents a vowel. A Bengali-English machine transliteration system has been developed based on the proposed models. The system has been trained to transliterate person names from Bengali to English. It uses the linguistic knowledge of possible conjuncts and diphthongs in Bengali and their equivalents in English. The system has been evaluated and it has been observed that the modified joint source-channel model performs best with a Word Agreement Ratio of 69.3% and a Transliteration Unit Agreement Ratio of 89.8%.
14
Kevin Knight, Jonathan Graehl · 1997 · 471 citations · Full text
Machine transliteration of names in Arabic text
Yaser Al-Onaizan, Kevin Knight · 2002 · 160 citations · Full text