A Toolbox for Construction and Analysis of Speech Datasets
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg
Abstract
Automatic Speech Recognition and Text-to-Speech systems are primarily trained in a supervised fashion and require high-quality, accurately labeled speech datasets. In this work, we examine common problems with speech data and introduce a toolbox for the construction and interactive error analysis of speech datasets. The construction tool is based on K{\"u}rzinger et al. work, and, to the best of our knowledge, the dataset exploration tool is the world's first open-source tool of this kind. We demonstrate how to apply these tools to create a Russian speech dataset and analyze existing speech datasets (Multilingual LibriSpeech, Mozilla Common Voice). The tools are open sourced as a part of the NeMo framework.
BibTeX
@inproceedings{
bakhturina2021a,
title={A Toolbox for Construction and Analysis of Speech Datasets},
author={Evelina Bakhturina and Vitaly Lavrukhin and Boris Ginsburg},
booktitle={Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)},
year={2021},
url={https://openreview.net/forum?id=oJ0oHQtAld}
}