Overcoming the Rare Word Problem for Low-Resource Language Pairs in Neural Machine Translation

Abstract

Among the six challenges of neural machine translation (NMT) coined by (Koehnand Knowles, 2017), rare-word problem is considered the most severe one,especially in translation of low-resource languages. In this paper, we proposethree solutions to address the rare words in neural machine translationsystems. First, we enhance source context to predict the target words byconnecting directly the source embeddings to the output of the attentioncomponent in NMT. Second, we propose an algorithm to learn morphology ofunknown words for English in supervised way in order to minimize the adverseeffect of rare-word problem. Finally, we exploit synonymous relation from theWordNet to overcome out-of-vocabulary (OOV) problem of NMT. We evaluate ourapproaches on two low-resource language pairs: English-Vietnamese andJapanese-Vietnamese. In our experiments, we have achieved significantimprovements of up to roughly +1.0 BLEU points in both language pairs.

Quick Read (beta)

loading the full paper ...