2019 · 110 citations · 34 references
Low‑resource language pairs suffer from inadequate parallel data, impairing adequacy and fluency, and data augmentation with abundant monolingual data is an effective remedy. The paper proposes a general data‑augmentation framework for low‑resource machine translation that leverages target‑side monolingual data and pivots through a related high‑resource language. The method uses a two‑step pivoting process that first injects low‑resource words into high‑resource sentences via an induced bilingual dictionary, then refines the resulting data with a modified unsupervised machine‑translation framework. Experiments on four low‑resource datasets show that, under extreme low‑resource settings, the proposed augmentation improves translation quality by 1.5 to 8 BLEU points over supervised back‑translation baselines.
Low-resource language pairs with a paucity of parallel data pose challenges for machine translation in terms of both adequacy and fluency. Data augmentation utilizing a large amount of monolingual data is regarded as an effective way to alleviate the problem. In this paper, we propose a general framework of data augmentation for low-resource machine translation not only using target-side monolingual data, but also by pivoting through a related high-resource language. Specifically, we experiment with a two-step pivoting method to convert high-resource data to the low-resource language, making best use of available resources to better approximate the true distribution of the low-resource language. First, we inject low-resource words into high-resource sentences through an induced bilingual dictionary. Second, we further edit the high-resource data injected with low-resource words using a modified unsupervised machine translation framework. Extensive experiments on four low-resource datasets show that under extreme low-resource settings, our data augmentation techniques improve translation quality by up to 1.5 to 8 BLEU points compared to supervised back-translation baselines.
34
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
A Generalized Solution of the Orthogonal Procrustes Problem
Peter H. Schönemann · Psychometrika · 1966 · 2K citations