Towards stable AI systems for Evaluating Arabic Pronunciations

  • 2025-08-27 05:49:15
  • Hadi Zaatiti, Hatem Hajri, Osama Abdullah, Nader Masmoudi
  • 0

Abstract

Modern Arabic ASR systems such as wav2vec 2.0 excel at word- andsentence-level transcription, yet struggle to classify isolated letters. Inthis study, we show that this phoneme-level task, crucial for languagelearning, speech therapy, and phonetic research, is challenging becauseisolated letters lack co-articulatory cues, provide no lexical context, andlast only a few hundred milliseconds. Recogniser systems must therefore relysolely on variable acoustic cues, a difficulty heightened by Arabic's emphatic(pharyngealized) consonants and other sounds with no close analogues in manylanguages. This study introduces a diverse, diacritised corpus of isolatedArabic letters and demonstrates that state-of-the-art wav2vec 2.0 modelsachieve only 35% accuracy on it. Training a lightweight neural network onwav2vec embeddings raises performance to 65%. However, adding a small amplitudeperturbation (epsilon = 0.05) cuts accuracy to 32%. To restore robustness, weapply adversarial training, limiting the noisy-speech drop to 9% whilepreserving clean-speech accuracy. We detail the corpus, training pipeline, andevaluation protocol, and release, on demand, data and code for reproducibility.Finally, we outline future work extending these methods to word- andsentence-level frameworks, where precise letter pronunciation remains critical.

 

Quick Read (beta)

loading the full paper ...