← Search

R. J. Skerry-Ryan

6 accepted papers

2021

Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis

ICASSP 2021accepted

We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs. The architecture extends the Tacotron model by incorporating a normalizing flow into the autoregressive decoder loop. Output waveforms are modeled as a sequence of non-overlapping fixed-leng…

Cited by 0SourceScholar
2020

Location-Relative Attention Mechanisms for Robust Long-Form Speech Synthesis

ICASSP 2020accepted

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show that these failures can be addressed using simple location-relative attention m…

Cited by 0SourceScholar
2019

Semi-supervised Training for Improving Data Efficiency in End-to-end Speech Synthesis

ICASSP 2019accepted

Although end-to-end text-to-speech (TTS) models such as Tacotron have shown excellent results, they typically require a sizable set of high-quality <;text, audio> pairs for training, which are expensive to collect. In this paper, we propose a semi-supervised training framework to improve the data ef…

Cited by 0SourceScholar
2018

Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

ICASSP 2018accepted

This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a voc…

Cited by 0SourceScholar