NeurIPS 2021poster66 citations

Multilingual Spoken Words Corpus

Mark Mazumder, Sharad Chitlangia, Colby Banbury, Yiping Kang, Juan Manuel Ciro, Keith Achorn, Daniel Galvez, Mark Sabini

Abstract

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 340,000 keywords, totaling 23.4 million 1-second spoken examples (over 6,000 hours). The dataset has many use cases, ranging from voice-enabled consumer devices to call center automation. We generate this dataset by applying forced alignment on crowd-sourced sentence-level audio to produce per-word timing estimates for extraction. All alignments are included in the dataset. We provide a detailed analysis of the contents of the data and contribute methods for detecting potential outliers. We report baseline accuracy metrics on keyword spotting models trained from our dataset compared to models trained on a manually-recorded keyword dataset. We conclude with our plans for dataset maintenance, updates, and open-sourced code.

keyword spottingspeech recognitionlow resource languages
BibTeX
@inproceedings{
mazumder2021multilingual,
title={Multilingual Spoken Words Corpus},
author={Mark Mazumder and Sharad Chitlangia and Colby Banbury and Yiping Kang and Juan Manuel Ciro and Keith Achorn and Daniel Galvez and Mark Sabini and Peter Mattson and David Kanter and Greg Diamos and Pete Warden and Josh Meyer and Vijay Janapa Reddi},
booktitle={Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)},
year={2021},
url={https://openreview.net/forum?id=c20jiJ5K2H}
}