ACL 2024findings8 citations

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

Tahir Javed, Janki Nawale, Eldho George, Sakshi Joshi, Kaushal Bhogale, Deovrat Mehendale, Ishvinder Sethi, Aparna Ananthanarayanan

Abstract

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a median of 73 hours per language. Through this paper, we share our journey of capturing the cultural, linguistic and demographic diversity of India to create a one-of-its-kind inclusive and representative dataset. More specifically, we share an open-source blueprint for data collection at scale comprising of standardised protocols, centralised tools, a repository of engaging questions, prompts and conversation scenarios spanning multiple domains and topics of interest, quality control mechanisms, comprehensive transcription guidelines and transcription tools. We hope that this open source blueprint will serve as a comprehensive starter kit for data collection efforts in other multilingual regions of the world. Using INDICVOICES, we build IndicASR, the first ASR model to support all the 22 languages listed in the 8th schedule of the Constitution of India.

BibTeX
@inproceedings{javed-etal-2024-indicvoices,
    title = "{I}ndic{V}oices: Towards building an Inclusive Multilingual Speech Dataset for {I}ndian Languages",
    author = "Javed, Tahir  and
      Nawale, Janki  and
      George, Eldho  and
      Joshi, Sakshi  and
      Bhogale, Kaushal  and
      Mehendale, Deovrat  and
      Sethi, Ishvinder  and
      Ananthanarayanan, Aparna  and
      Faquih, Hafsah  and
      Palit, Pratiti  and
      Ravishankar, Sneha  and
      Sukumaran, Saranya  and
      Panchagnula, Tripura  and
      Murali, Sunjay  and
      Gandhi, Kunal  and
      R, Ambujavalli  and
      M, Manickam  and
      Vaijayanthi, C  and
      Karunganni, Krishnan  and
      Kumar, Pratyush  and
      Khapra, Mitesh",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-acl.639/",
    doi = "10.18653/v1/2024.findings-acl.639",
    pages = "10740--10782"
}
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages · ACL 2024