Speech Re-Painting for Robust ASR
Kyle Kastner, Gary Wang, Isaac Elias, Takaaki Saeki, Pedro Moreno Mengibar, Françoise Beaufays, Andrew Rosenberg, Bhuvana Ramabhadran
Abstract
Synthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synthetic data augmentation can be used to approximate the variability of natural speech, however, not all aspects of variation are equally relevant for augmentation. In this work, we introduce speech re-painting, a method for in-context augmented synthesis, using target training datasets to generate new utterances guided by speech and text on the fly in a zero-shot manner. We evaluate this technique using downstream ASR word error rate (WER) using the VCTK and LibriSpeech datasets. These represent unique speaker and lexical challenges that are addressed by re-painting, realizing a reduction of WER more than 50% in particular settings.
BibTeX
@inproceedings{icassp2025_speechrepainting,
title = {Speech Re-Painting for Robust ASR},
author = {Kyle Kastner and Gary Wang and Isaac Elias and Takaaki Saeki and Pedro Moreno Mengibar and Françoise Beaufays and Andrew Rosenberg and Bhuvana Ramabhadran},
booktitle = {ICASSP 2025},
year = {2025}
}