The TALP-UPC System for the WMT Similar Language Task: Statistical vs Neural Machine Translation

Abstract

Although the problem of similar language translation has been an area ofresearch interest for many years, yet it is still far from being solved. Inthis paper, we study the performance of two popular approaches: statistical andneural. We conclude that both methods yield similar results; however, theperformance varies depending on the language pair. While the statisticalapproach outperforms the neural one by a difference of 6 BLEU points for theSpanish-Portuguese language pair, the proposed neural model surpasses thestatistical one by a difference of 2 BLEU points for Czech-Polish. In theformer case, the language similarity (based on perplexity) is much higher thanin the latter case. Additionally, we report negative results for the systemcombination with back-translation. Our TALP-UPC system submission won 1st placefor Czech-to-Polish and 2nd place for Spanish-to-Portuguese in the officialevaluation of the 1st WMT Similar Language Translation task.

Quick Read (beta)

loading the full paper ...