ICLR 2018poster586 citations

Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning

Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, John Miller

Abstract

We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural speech synthesis systems in naturalness while training an order of magnitude faster. We scale Deep Voice 3 to dataset sizes unprecedented for TTS, training on more than eight hundred hours of audio from over two thousand speakers. In addition, we identify common error modes of attention-based speech synthesis networks, demonstrate how to mitigate them, and compare several different waveform synthesis methods. We also describe how to scale inference to ten million queries per day on a single GPU server.

2000-Speaker Neural TTSMonotonic AttentionSpeech Synthesis
BibTeX
@inproceedings{
ping2018deep,
title={Deep Voice 3: 2000-Speaker Neural Text-to-Speech},
author={Wei Ping and Kainan Peng and Andrew Gibiansky and Sercan O. Arik and Ajay Kannan and Sharan Narang and Jonathan Raiman and John Miller},
booktitle={International Conference on Learning Representations},
year={2018},
url={https://openreview.net/forum?id=HJtEm4p6Z},
}
Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning · ICLR 2018