Abstract
Low-resource languages present unique challenges to (neural) machinetranslation. We discuss the case of Bambara, a Mande language for whichtraining data is scarce and requires significant amounts of pre-processing.More than the linguistic situation of Bambara itself, the socio-culturalcontext within which Bambara speakers live poses challenges for automatedprocessing of this language. In this paper, we present the first parallel dataset for machine translation of Bambara into and from English and French and thefirst benchmark results on machine translation to and from Bambara. We discusschallenges in working with low-resource languages and propose strategies tocope with data scarcity in low-resource machine translation (MT).
Quick Read (beta)
loading the full paper ...