← Search

Rob Clark

3 accepted papers

2022

Improving Phonetic Realizations in its by Using Phoneme-Aligned Graphemes

ICASSP 2022accepted

Most text-to-speech acoustic models, such as WaveNet, Tacotron, ClariNet, etc., use either a phoneme sequence or a letter sequence as the fundamental unit of speech. Although the letter (or grapheme) sequence closely matches the actual runtime input of the TTS system, it often fails to represent the…

Cited by 0SourceScholar
2019

CHiVE: Varying Prosody in Speech Synthesis with a Linguistically Driven Dynamic Hierarchical Conditional Variational Network

ICML 2019oral

The prosodic aspects of speech signals produced by current text-to-speech systems are typically averaged over training material, and as such lack the variety and liveliness found in natural speech. To avoid monotony and averaged prosody contours, it is desirable to have a way of modeling the variati…

Cited by 111SourcePDFScholar
2018

Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

ICML 2018oral

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing the desired prosody. We show that conditioning Tacotron on this learned embedding space results in synthesized audio that…

Cited by 749SourcePDFScholar