Replacing Human Audio with Synthetic Audio for on-Device Unspoken Punctuation Prediction
Daria Soboleva, Ondrej Skopek, Márius Sajgalík, Victor Carbune, Felix Weissenberger, Julia Proskurnia, Bogdan Prisacari, Daniel Valcarce
Abstract
We present a novel multi-modal unspoken punctuation prediction system for the English language which combines acoustic and text features. We demonstrate for the first time, that by relying exclusively on synthetic data generated using a prosody-aware text-to-speech system, we can outperform a model trained with expensive human audio recordings on the unspoken punctuation prediction problem. Our model architecture is well suited for on-device use. This is achieved by leveraging hash-based embeddings of automatic speech recognition text output in conjunction with acoustic features as input to a quasi-recurrent neural network, keeping the model size small and latency low.
BibTeX
@inproceedings{icassp2021_replacinghumanau,
title = {Replacing Human Audio with Synthetic Audio for on-Device Unspoken Punctuation Prediction},
author = {Daria Soboleva and Ondrej Skopek and Márius Sajgalík and Victor Carbune and Felix Weissenberger and Julia Proskurnia and Bogdan Prisacari and Daniel Valcarce and Justin Lu and Rohit Prabhavalkar and Balint Miklos},
booktitle = {ICASSP 2021},
year = {2021}
}