EngineeringMachine LearningMultilingual PretrainingCorpus LinguisticsSpeech RecognitionNatural Language ProcessingComputational LinguisticsLanguage StudiesReal-time LanguageRobust TranslationMachine TranslationComputer-assisted TranslationLinguisticsComputer ScienceNeural Machine TranslationAutomatic Speech RecognitionAsr ErrorsSpeech ProcessingSpeech Translation
In many practical applications, neural machine translation systems have to deal with the input from automatic speech recognition (ASR) systems which may contain a certain number of errors. This leads to two problems which degrade translation performance. One is the discrepancy between the training and testing data and the other is the translation error caused by the input errors may ruin the whole translation. In this paper, we propose a method to handle the two problems so as to generate robust translation to ASR errors. First, we simulate ASR errors in the training data so that the data distribution in the training and test is consistent. Second, we focus on ASR errors on homophone words and words with similar pronunciation and make use of their pronunciation information to help the translation model to recover from the input errors. Experiments on two Chinese-English data sets show that our method is more robust to input errors and can outperform the strong Transformer baseline significantly.
20
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
Kishore Papineni, Salim Roukos, Todd J. Ward et al. · 2001 · 20.9K citations · Full text
Natural Language Processing, Computer-assisted Translation, Engineering +10
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, Quoc V. Le · arXiv (Cornell University) · 2014 · 13.3K citations · Full text