AAAI 2024technical0 citations

Towards Building a Language-Independent Speech Scoring Assessment

Shreyansh Gupta, Abhishek Unnam, Kuldeep Yadav, Varun Aggarwal

Abstract

Automatic speech scoring is crucial in language learning, providing targeted feedback to language learners by assessing pronunciation, fluency, and other speech qualities. However, the scarcity of human-labeled data for languages beyond English poses a significant challenge in developing such systems. In this work, we propose a Language-Independent scoring approach to evaluate speech without relying on labeled data in the target language. We introduce a multilingual speech scoring system that leverages representations from the wav2vec 2.0 XLSR model and a force-alignment technique based on CTC-Segmentation to construct speech features. These features are used to train a machine learning model to predict pronunciation and fluency scores. We demonstrate the potential of our method by predicting expert ratings on a speech dataset spanning five languages - English, French, Spanish, German and Portuguese, and comparing its performance against Language-Specific models trained individually on each language, as well as a jointly-trained model on all languages. Results indicate that our approach shows promise as an initial step towards a universal language independent speech scoring.

BibTeX
@article{Gupta_Unnam_Yadav_Aggarwal_2024, title={Towards Building a Language-Independent Speech Scoring Assessment}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/30366}, DOI={10.1609/aaai.v38i21.30366}, abstractNote={Automatic speech scoring is crucial in language learning, providing targeted feedback to language learners by assessing pronunciation, fluency, and other speech qualities. However, the scarcity of human-labeled data for languages beyond English poses a significant challenge in developing such systems. In this work, we propose a Language-Independent scoring approach to evaluate speech without relying on labeled data in the target language. We introduce a multilingual speech scoring system that leverages representations from the wav2vec 2.0 XLSR model and a force-alignment technique based on CTC-Segmentation to construct speech features. These features are used to train a machine learning model to predict pronunciation and fluency scores. We demonstrate the potential of our method by predicting expert ratings on a speech dataset spanning five languages - English, French, Spanish, German and Portuguese, and comparing its performance against Language-Specific models trained individually on each language, as well as a jointly-trained model on all languages. Results indicate that our approach shows promise as an initial step towards a universal language independent speech scoring.}, number={21}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Gupta, Shreyansh and Unnam, Abhishek and Yadav, Kuldeep and Aggarwal, Varun}, year={2024}, month={Mar.}, pages={23200-23206} }