Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System

Abstract

In this work we study some of the challenges associated with scaling speakerrecognition systems to multiple languages. To the best of our knowledge, thisis the first study of speaker verification systems at the scale of 46languages. Training models for each of the many languages can be time andenergy demanding in addition to costly. Low resource languages presentadditional difficulties. The problem is framed from the perspective of using asmart speaker device with interactions consisting of a wake-up keyword(text-dependent) followed by a speech query (text-independent). We examine the use of a hybrid setup consisting of multilingualtext-dependent and text-independent components. Experimental evidence suggeststhat training on multiple languages can generalize to unseen varieties whilemaintaining performance on seen varieties. We also found that it can reducecomputational requirements for training models by an order of magnitude.Furthermore, during model inference on English data, we observe that leveraginga triage framework can reduce the number of calls to the more computationallyexpensive text-independent system by 73% (and reduce latency by 60%) whilemaintaining an EER no worse than the text-independent setup.

Quick Read (beta)

loading the full paper ...