Neural Language Codes for Multilingual Acoustic Models

Abstract

Multilingual Speech Recognition is one of the most costly AI problems,because each language (7,000+) and even different accents require their ownacoustic models to obtain best recognition performance. Even though they alluse the same phoneme symbols, each language and accent imposes its own coloringor "twang". Many adaptive approaches have been proposed, but they requirefurther training, additional data and generally are inferior to monolinguallytrained models. In this paper, we propose a different approach that uses alarge multilingual model that is \emph{modulated} by the codes generated by anancillary network that learns to code useful differences between the "twangs"or human language. We use Meta-Pi networks to have one network (the language code net) gate theactivity of neurons in another (the acoustic model nets). Our results show thatduring recognition multilingual Meta-Pi networks quickly adapt to the properlanguage coloring without retraining or new data, and perform better thanmonolingually trained networks. The model was evaluated by training acousticmodeling nets and modulating language code nets jointly and optimize them forbest recognition performance.

Quick Read (beta)

loading the full paper ...