← Search

Arnaud Joly

5 accepted papers

2024

Mapache: Masked Parallel Transformer for Advanced Speech Editing and Synthesis

ICASSP 2024accepted

Recent advancements in Generative AI, such as scaled Transformer large language models (LLM) and diffusion decoders, have revolutionized speech synthesis. With speech encompassing the complexities of natural language and audio dimensionality, many recent models have relied on autoregressive modeling…

Cited by 0SourceScholar
2022

Distribution Augmentation for Low-Resource Expressive Text-To-Speech

ICASSP 2022accepted

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data. Our goal is to in-crease diversity of text conditionings available during training. This helps to reduce overfitting, e…

Cited by 0SourceScholar
2021

Camp: A Two-Stage Approach to Modelling Prosody in Context

ICASSP 2021accepted

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background n…

Cited by 33SourceScholar
2021

Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

ICASSP 2021accepted

In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn a prosodic distribution at the sentence level from mel-spectrograms available during training. In Stage II, we propose…

Cited by 0SourceScholar