Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon

  • 2021-02-15 09:27:54
  • Hadi Veisi, Hawre Hosseini, Mohammad Mohammadamini, Wirya Fathy, Aso Mahmudi
  • 0


In this paper, we introduce the first large vocabulary speech recognitionsystem (LVSR) for the Central Kurdish language, named Jira. The Kurdishlanguage is an Indo-European language spoken by more than 30 million people inseveral countries, but due to the lack of speech and text resources, there isno speech recognition system for this language. To fill this gap, we introducethe first speech corpus and pronunciation lexicon for the Kurdish language.Regarding speech corpus, we designed a sentence collection in which the ratioof di-phones in the collection resembles the real data of the Central Kurdishlanguage. The designed sentences are uttered by 576 speakers in a controlledenvironment with noise-free microphones (called AsoSoft Speech-Office) and inTelegram social network environment using mobile phones (denoted as AsoSoftSpeech-Crowdsourcing), resulted in 43.68 hours of speech. Besides, a test setincluding 11 different document topics is designed and recorded in twocorresponding speech conditions (i.e., Office and Crowdsourcing). Furthermore,a 60K pronunciation lexicon is prepared in this research in which we facedseveral challenges and proposed solutions for them. The Kurdish language hasseveral dialects and sub-dialects that results in many lexical variations. Ourmethods for script standardization of lexical variations and automaticpronunciation of the lexicon tokens are presented in detail. To setup therecognition engine, we used the Kaldi toolkit. A statistical tri-gram languagemodel that is extracted from the AsoSoft text corpus is used in the system.Several standard recipes including HMM-based models (i.e., mono, tri1, tr2,tri2, tri3), SGMM, and DNN methods are used to generate the acoustic model.These methods are trained with AsoSoft Speech-Office and AsoSoftSpeech-Crowdsourcing and a combination of them. The best performance achievedby the SGMM acoustic model which results in 13.9% of the average word errorrate (on different document topics) and 4.9% for the general topic.


Quick Read (beta)

loading the full paper ...