← Search

Jaime Lorenzo-Trueba

11 accepted papers

2022

Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation

ICASSP 2022accepted

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting conversational expressive data from different speakers. Our goal is to build a…

Cited by 0SourceScholar
2022

Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module

ICASSP 2022accepted

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and intelligibility degradations, making training low-resource TTS system…

Cited by 31SourceScholar
2021

Camp: A Two-Stage Approach to Modelling Prosody in Context

ICASSP 2021accepted

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background n…

Cited by 33SourceScholar
2021

Low-Resource Expressive Text-To-Speech Using Data Augmentation

ICASSP 2021accepted

While recent neural text-to-speech (TTS) systems perform remarkably well, they typically require a substantial amount of recordings from the target speaker reading in the desired speaking style. In this work, we present a novel 3-step methodology to circumvent the costly operation of recording large…

Cited by 0SourceScholar
2021

Mispronunciation Detection in Non-Native (L2) English with Uncertainty Modeling

ICASSP 2021accepted

A common approach to the automatic detection of mispronunciation in language learning is to recognize the phonemes produced by a student and compare it to the expected pronunciation of a native speaker. This approach makes two simplifying assumptions: a) phonemes can be recognized from speech with h…

Cited by 0SourceScholar
2021

Proteno: Text Normalization with Limited Data for Fast Deployment in Text to Speech Systems

NAACL 2021industry

Developing Text Normalization (TN) systems for Text-to-Speech (TTS) on new languages is hard. We propose a novel architecture to facilitate it for multiple languages while using data less than 3% of the size of the data used by the state of the art results on English. We treat TN as a sequence class…

2020

Using Vaes and Normalizing Flows for One-Shot Text-To-Speech Synthesis of Expressive Speech

ICASSP 2020accepted

We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence based system with a Variational AutoEncoder (VAE) and a Househol…

Cited by 33SourceScholar
2019

Effect of Data Reduction on Sequence-to-sequence Neural TTS

ICASSP 2019accepted

Recent speech synthesis systems based on sampling from autoregressive neural network models can generate speech almost indistinguishable from human recordings. However, these models require large amounts of data. This paper shows that the lack of data from one speaker can be compensated with data fr…

Cited by 63SourceScholar
2018

A Comparison of Recent Waveform Generation and Acoustic Modeling Methods for Neural-Network-Based Speech Synthesis

ICASSP 2018accepted

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in wh…

Cited by 0SourceScholar
2018

Cyborg Speech: Deep Multilingual Speech Synthesis for Generating Segmental Foreign Accent with Natural Prosody

ICASSP 2018accepted

We describe a new application of deep-learning-based speech synthesis, namely multilingual speech synthesis for generating controllable foreign accent. Specifically, we train a DBLSTM-based acoustic model on non-accented multilingual speech recordings from a speaker native in several languages. By c…

Cited by 0SourceScholar
2018

High-Quality Nonparallel Voice Conversion Based on Cycle-Consistent Adversarial Network

ICASSP 2018accepted

Although voice conversion (VC) algorithms have achieved remarkable success along with the development of machine learning, superior performance is still difficult to achieve when using nonparallel data. In this paper, we propose using a cycle-consistent adversarial network (CycleGAN) for nonparallel…

Cited by 0SourceScholar