Getting Gender Right in Neural Machine Translation

  • 2019-09-11 14:44:27
  • Eva Vanmassenhove, Christian Hardmeier, Andy Way
  • 2

Abstract

Speakers of different languages must attend to and encode strikinglydifferent aspects of the world in order to use their language correctly (Sapir,1921; Slobin, 1996). One such difference is related to the way gender isexpressed in a language. Saying "I am happy" in English, does not encode anyadditional knowledge of the speaker that uttered the sentence. However, manyother languages do have grammatical gender systems and so such knowledge wouldbe encoded. In order to correctly translate such a sentence into, say, French,the inherent gender information needs to be retained/recovered. The samesentence would become either "Je suis heureux", for a male speaker or "Je suisheureuse" for a female one. Apart from morphological agreement, demographicfactors (gender, age, etc.) also influence our use of language in terms of wordchoices or even on the level of syntactic constructions (Tannen, 1991;Pennebaker et al., 2003). We integrate gender information into NMT systems. Ourcontribution is two-fold: (1) the compilation of large datasets with speakerinformation for 20 language pairs, and (2) a simple set of experiments thatincorporate gender information into NMT for multiple language pairs. Ourexperiments show that adding a gender feature to an NMT system significantlyimproves the translation quality for some language pairs.

 

Quick Read (beta)

loading the full paper ...