← Search

Éva Székely

6 accepted papers

2024

Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching

ICASSP 2024accepted

We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output quality in fewer synthesis steps than models trained using score matching. Carefu…

Cited by 252SourceScholar
2024

Unified Speech and Gesture Synthesis Using Flow Matching

ICASSP 2024accepted

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures. This paper presents a novel, unified architecture for jointl…

Cited by 7SourceScholar
2023

Prosody-Controllable Spontaneous TTS with Neural HMMS

ICASSP 2023accepted

Spontaneous speech has many affective and pragmatic functions that are interesting and challenging to model in TTS. However, the presence of reduced articulation, fillers, repetitions, and other disfluencies in spontaneous speech make the text and acoustics less aligned than in read speech, which is…

Cited by 0SourceScholar
2022

Neural HMMS Are All You Need (For High-Quality Attention-Free TTS)

ICASSP 2022accepted

Neural sequence-to-sequence TTS has achieved significantly better output quality than statistical speech synthesis using HMMs. However, neural TTS is generally not probabilistic and uses non-monotonic attention. Attention failures increase training time and can make synthesis babble incoherently. Th…

Cited by 0SourceScholar
2020

Breathing and Speech Planning in Spontaneous Speech Synthesis

ICASSP 2020accepted

Breathing and speech planning in spontaneous speech are coordinated processes, often exhibiting disfluent patterns. While synthetic speech is not subject to respiratory needs, integrating breath into synthesis has advantages for naturalness and recall. At the same time, a synthetic voice reproducing…

Cited by 36SourceScholar
2019

Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath Detector

ICASSP 2019accepted

This paper considers utilising breaths to create improved spontaneous-speech corpora for conversational text-to-speech from found audio recordings such as dialogue podcasts. Breaths are of interest since they relate to prosody and speech planning and are independent of language and transcription. Sp…

Cited by 44SourceScholar