← Search

David Kao

3 accepted papers

2020

Location-Relative Attention Mechanisms for Robust Long-Form Speech Synthesis

ICASSP 2020accepted

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show that these failures can be addressed using simple location-relative attention m…

Cited by 0SourceScholar
2020

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

ICLR 2020poster

We present a novel generative model that combines state-of-the-art neural text- to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial supervision to some of the latent variables, we are able to force them to take on consistent and interpretable purposes, whi…

Cited by 61SourcecodeScholar