Multilingual Neural Machine Translation:Can Linguistic Hierarchies Help?

  • 2021-10-15 02:31:48
  • Fahimeh Saleh, Wray Buntine, Gholamreza Haffari, Lan Du
  • 1

Abstract

Multilingual Neural Machine Translation (MNMT) trains a single NMT model thatsupports translation between multiple languages, rather than training separatemodels for different languages. Learning a single model can enhance thelow-resource translation by leveraging data from multiple languages. However,the performance of an MNMT model is highly dependent on the type of languagesused in training, as transferring knowledge from a diverse set of languagesdegrades the translation performance due to negative transfer. In this paper,we propose a Hierarchical Knowledge Distillation (HKD) approach for MNMT whichcapitalises on language groups generated according to typological features andphylogeny of languages to overcome the issue of negative transfer. HKDgenerates a set of multilingual teacher-assistant models via a selectiveknowledge distillation mechanism based on the language groups, and then distilsthe ultimate multilingual model from those assistants in an adaptive way.Experimental results derived from the TED dataset with 53 languages demonstratethe effectiveness of our approach in avoiding the negative transfer effect inMNMT, leading to an improved translation performance (about 1 BLEU score onaverage) compared to strong baselines.

 

Quick Read (beta)

loading the full paper ...