← Search

Zack Hodari

4 accepted papers

2024

Controllable Prosody Generation with Partial Inputs

ICASSP 2024accepted

We address the problem of human-in-the-loop control for generating prosody in the context of text-to-speech synthesis. Controlling prosody is challenging because existing generative models lack an efficient interface through which users can modify the output quickly and precisely. To solve this, we…

Cited by 0SourceScholar
2023

Ensemble Prosody Prediction For Expressive Speech Synthesis

ICASSP 2023accepted

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Yet, in evaluations it is generally found that no single model is preferred for al…

Cited by 4SourceScholar
2021

Camp: A Two-Stage Approach to Modelling Prosody in Context

ICASSP 2021accepted

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background n…

Cited by 33SourceScholar
2021

Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

ICASSP 2021accepted

In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn a prosodic distribution at the sentence level from mel-spectrograms available during training. In Stage II, we propose…

Cited by 0SourceScholar