← Search

Deovrat Mehendale

3 accepted papers

2025

Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages

ICASSP 2025accepted

Lack of large-scale pretraining data for low resource languages from the Indian sub-continent, leads to their underrepresentation in existing massively multilingual models. In this work, we address this gap by proposing a framework to create large raw audio datasets for such under-represented langua…

Cited by 0SourceScholar
2024

IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS

NeurIPS 2024poster

Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack of high-quality, manually subtitled data on platforms like LibriVox or YouTube.…

2024

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

ACL 2024findings

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a…