← Search

Sakshi Joshi

2 accepted papers

2025

Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages

ICASSP 2025accepted

Lack of large-scale pretraining data for low resource languages from the Indian sub-continent, leads to their underrepresentation in existing massively multilingual models. In this work, we address this gap by proposing a framework to create large raw audio datasets for such under-represented langua…

Cited by 0SourceScholar
2024

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

ACL 2024findings

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a…