← Search

Gustav Eje Henter

15 accepted papers

2024

Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching

ICASSP 2024accepted

We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output quality in fewer synthesis steps than models trained using score matching. Carefu…

Cited by 252SourceScholar
2024

Unified Speech and Gesture Synthesis Using Flow Matching

ICASSP 2024accepted

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures. This paper presents a novel, unified architecture for jointl…

Cited by 7SourceScholar
2023

A Processing Framework to Access Large Quantities of Whispered Speech Found in ASMR

ICASSP 2023accepted

Whispering is a ubiquitous mode of communication that humans use daily. Despite this, whispered speech has been poorly served by existing speech technology due to a shortage of resources and processing methodology. To remedy this, this paper provides a processing framework that enables access to lar…

Cited by 0SourceScholar
2023

Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing

ICASSP 2023accepted

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation.A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from…

Cited by 0SourceScholar
2023

Prosody-Controllable Spontaneous TTS with Neural HMMS

ICASSP 2023accepted

Spontaneous speech has many affective and pragmatic functions that are interesting and challenging to model in TTS. However, the presence of reduced articulation, fillers, repetitions, and other disfluencies in spontaneous speech make the text and acoustics less aligned than in read speech, which is…

Cited by 0SourceScholar
2022

Neural HMMS Are All You Need (For High-Quality Attention-Free TTS)

ICASSP 2022accepted

Neural sequence-to-sequence TTS has achieved significantly better output quality than statistical speech synthesis using HMMs. However, neural TTS is generally not probabilistic and uses non-monotonic attention. Attention failures increase training time and can make synthesis babble incoherently. Th…

Cited by 0SourceScholar
2022

Wavebender GAN: An Architecture for Phonetically Meaningful Speech Manipulation

ICASSP 2022accepted

Deep learning has revolutionised synthetic speech quality. However, it has thus far delivered little value to the speech science community. The new methods do not meet the controllability demands that practitioners in this area require e.g.: in listening tests with manipulated speech stimuli. Instea…

Cited by 10SourceScholar
2021

The Case for Translation-Invariant Self-Attention in Transformer-Based Language Models

ACL 2021short

Mechanisms for encoding positional information are central for transformer-based language models. In this paper, we analyze the position embeddings of existing language models, finding strong evidence of translation invariance, both for the embeddings themselves and for their effect on self-attentio…

2020

Breathing and Speech Planning in Spontaneous Speech Synthesis

ICASSP 2020accepted

Breathing and speech planning in spontaneous speech are coordinated processes, often exhibiting disfluent patterns. While synthetic speech is not subject to respiratory needs, integrating breath into synthesis has advantages for naturalness and recall. At the same time, a synthetic voice reproducing…

Cited by 36SourceScholar
2019

Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath Detector

ICASSP 2019accepted

This paper considers utilising breaths to create improved spontaneous-speech corpora for conversational text-to-speech from found audio recordings such as dialogue podcasts. Breaths are of interest since they relate to prosody and speech planning and are independent of language and transcription. Sp…

Cited by 44SourceScholar
2018

Cyborg Speech: Deep Multilingual Speech Synthesis for Generating Segmental Foreign Accent with Natural Prosody

ICASSP 2018accepted

We describe a new application of deep-learning-based speech synthesis, namely multilingual speech synthesis for generating controllable foreign accent. Specifically, we train a DBLSTM-based acoustic model on non-accented multilingual speech recordings from a speaker native in several languages. By c…

Cited by 0SourceScholar
2017

Adapting and controlling DNN-based speech synthesis using input codes

ICASSP 2017accepted

Methods for adapting and controlling the characteristics of output speech are important topics in speech synthesis. In this work, we investigated the performance of DNN-based text-to-speech systems that in parallel to conventional text input also take speaker, gender, and age codes as inputs, in ord…

Cited by 0SourceScholar
2016

From HMMS to DNNS: Where do the improvements come from?

ICASSP 2016accepted

Deep neural networks (DNNs) have recently been the focus of much text-to-speech research as a replacement for decision trees and hidden Markov models (HMMs) in statistical parametric synthesis systems. Performance improvements have been reported; however, the configuration of systems evaluated makes…

Cited by 0SourceScholar
2016

Robust TTS duration modelling using DNNS

ICASSP 2016accepted

Accurate modelling and prediction of speech-sound durations is an important component in generating more natural synthetic speech. Deep neural networks (DNNs) offer a powerful modelling paradigm, and large, found corpora of natural and expressive speech are easy to acquire for training them. Unfortu…

Cited by 0SourceScholar
2016

Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis

ICASSP 2016accepted

Forced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do n…

Cited by 14SourceScholar