Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASR
Amartyaveer, Saurabh Kumar, Sumit Sharma, Sathvik Udupa, Sandhya Badiger, Abhayjeet Singh, Deekshitha G, Jesuraja Bandekar
Abstract
Dialect identification (DID) addresses the challenge of recog-nizing regional variations within a language. The current deep learning approaches focus on audio-only, text-only, or multi-task setups combining automatic speech recognition (ASR) with DID. This work introduces a novel multimodal architecture that leverages speech and text features to enhance DID performance. Our method integrates ASR-generated speech representations with text embeddings derived from ASR hypotheses using a RoBERTa-based encoder. Additionally, we perform a layer-wise analysis of the IndicWav2Vec model to identify the layers most effective for extracting dialectal features. We evaluate our approach on a subset of the RESPIN dataset featuring eight Indian languages and 33 dialects. Experimental results show that our proposed multimodal DID system achieves an average DID accuracy of 79.81%, consistently outperforming baseline methods. This study is the first to analyse comprehensively DID in Indian languages, providing new insights into their dialectal diversity.
BibTeX
@inproceedings{icassp2025_improvingdialect,
title = {Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASR},
author = {Amartyaveer and Saurabh Kumar and Sumit Sharma and Sathvik Udupa and Sandhya Badiger and Abhayjeet Singh and Deekshitha G and Jesuraja Bandekar and Savitha Murthy and Prasanta Kumar Ghosh},
booktitle = {ICASSP 2025},
year = {2025}
}