ICASSP 2021accepted0 citations

Exploring the application of synthetic audio in training keyword spotters

Andrew Werchniak, Roberto Barra-Chicote, Yuriy Mishchenko, Jasha Droppo, Jeff Condal, Peng Liu, Anish Shah

Abstract

The study of keyword spotting, a subfield within the broader field of speech recognition that centers around identifying individual keywords in speech audio, has gained particular importance in recent years with the rise of personal voice assistants such as Alexa. As voice assistants aim to rapidly expand to support new languages, keywords, and use cases, stakeholders face the issue of limited training data for these unseen scenarios. This paper details some initial exploration into the application of Text-To-Speech (TTS) audio as a "helper" tool for training keyword spotters in these low-resource scenarios. In the experiments studied in this paper, the careful mixing of TTS audio with human speech audio during training led to a reduction of over 11% in the detection-error-tradeoff (DET) area under the curve (AUC) metric.

BibTeX
@inproceedings{icassp2021_exploringtheappl,
  title = {Exploring the application of synthetic audio in training keyword spotters},
  author = {Andrew Werchniak and Roberto Barra-Chicote and Yuriy Mishchenko and Jasha Droppo and Jeff Condal and Peng Liu and Anish Shah},
  booktitle = {ICASSP 2021},
  year = {2021}
}