← Search

Thomas Drugman

6 accepted papers

2022

Distribution Augmentation for Low-Resource Expressive Text-To-Speech

ICASSP 2022accepted

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data. Our goal is to in-crease diversity of text conditionings available during training. This helps to reduce overfitting, e…

Cited by 0SourceScholar
2021

Camp: A Two-Stage Approach to Modelling Prosody in Context

ICASSP 2021accepted

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background n…

Cited by 33SourceScholar
2021

Mispronunciation Detection in Non-Native (L2) English with Uncertainty Modeling

ICASSP 2021accepted

A common approach to the automatic detection of mispronunciation in language learning is to recognize the phonemes produced by a student and compare it to the expected pronunciation of a native speaker. This approach makes two simplifying assumptions: a) phonemes can be recognized from speech with h…

Cited by 0SourceScholar
2021

Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

ICASSP 2021accepted

In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn a prosodic distribution at the sentence level from mel-spectrograms available during training. In Stage II, we propose…

Cited by 0SourceScholar
2019

Effect of Data Reduction on Sequence-to-sequence Neural TTS

ICASSP 2019accepted

Recent speech synthesis systems based on sampling from autoregressive neural network models can generate speech almost indistinguishable from human recordings. However, these models require large amounts of data. This paper shows that the lack of data from one speaker can be compensated with data fr…

Cited by 63SourceScholar
2015

Robust excitation-based features for Automatic Speech Recognition

ICASSP 2015accepted

In this paper we investigate the use of noise-robust features characterizing the speech excitation signal as complementary features to the usually considered vocal tract based features for Automatic Speech Recognition (ASR). The proposed Excitation-based Features (EBF) are tested in a state-of-the-a…

Cited by 0SourceScholar