← Search

Tahir Javed

5 accepted papers

2025

Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages

ICASSP 2025accepted

Lack of large-scale pretraining data for low resource languages from the Indian sub-continent, leads to their underrepresentation in existing massively multilingual models. In this work, we address this gap by proposing a framework to create large raw audio datasets for such under-represented langua…

Cited by 0SourceScholar
2024

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

ACL 2024findings

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a…

2023

Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages

ICASSP 2023accepted

Collecting labelled datasets for speech recognition systems for low-resource languages on a diverse set of domains and speakers is expensive. In this work, we demonstrate an inexpensive and effective alternative by "mining" text and audio pairs for Indian languages from public sources, specifically…

Cited by 0SourceScholar
2023

IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian Languages

AAAI 2023technical

A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. Th…

2022

Towards Building ASR Systems for the Next Billion Users

AAAI 2022technical

Recent methods in speech and language technology pretrain very large models which are fine-tuned for specific tasks. However, the benefits of such large models are often limited to a few resource rich languages of the world. In this work, we make multiple contributions towards building ASR systems f…