← Search

Thomas Merritt

8 accepted papers

2023

AE-Flow: Autoencoder Normalizing Flow

ICASSP 2023accepted

Recently normalizing flows have been gaining traction in text-to-speech (TTS) and voice conversion (VC) due to their state-of-the-art (SOTA) performance. Normalizing flows are unsupervised generative models. In this paper, we introduce supervision to the training process of normalizing flows, withou…

Cited by 0SourceScholar
2022

Text-Free Non-Parallel Many-To-Many Voice Conversion Using Normalising Flow

ICASSP 2022accepted

Non-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all other information from the source speech is retained is a large challenge. This is particularly challenging in the scen…

Cited by 0SourceScholar
2021

Camp: A Two-Stage Approach to Modelling Prosody in Context

ICASSP 2021accepted

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background n…

Cited by 33SourceScholar
2021

Low-Resource Expressive Text-To-Speech Using Data Augmentation

ICASSP 2021accepted

While recent neural text-to-speech (TTS) systems perform remarkably well, they typically require a substantial amount of recordings from the target speaker reading in the desired speaking style. In this work, we present a novel 3-step methodology to circumvent the costly operation of recording large…

Cited by 0SourceScholar
2019

Effect of Data Reduction on Sequence-to-sequence Neural TTS

ICASSP 2019accepted

Recent speech synthesis systems based on sampling from autoregressive neural network models can generate speech almost indistinguishable from human recordings. However, these models require large amounts of data. This paper shows that the lack of data from one speaker can be compensated with data fr…

Cited by 63SourceScholar
2016

Deep neural network-guided unit selection synthesis

ICASSP 2016accepted

Vocoding of speech is a standard part of statistical parametric speech synthesis systems. It imposes an upper bound of the naturalness that can possibly be achieved. Hybrid systems using parametric models to guide the selection of natural speech units can combine the benefits of robust statistical m…

Cited by 0SourceScholar
2016

From HMMS to DNNS: Where do the improvements come from?

ICASSP 2016accepted

Deep neural networks (DNNs) have recently been the focus of much text-to-speech research as a replacement for decision trees and hidden Markov models (HMMs) in statistical parametric synthesis systems. Performance improvements have been reported; however, the configuration of systems evaluated makes…

Cited by 0SourceScholar
2015

Attributing modelling errors in HMM synthesis by stepping gradually from natural to modelled speech

ICASSP 2015accepted

Even the best statistical parametric speech synthesis systems do not achieve the naturalness of good unit selection. We investigated possible causes of this. By constructing speech signals that lie in between natural speech and the output from a complete HMM synthesis system, we investigated various…

Cited by 20SourceScholar