Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages
Lack of large-scale pretraining data for low resource languages from the Indian sub-continent, leads to their underrepresentation in existing massively multilingual models. In this work, we address this gap by proposing a framework to create large raw audio datasets for such under-represented langua…