← Search

Frank Zalkow

3 accepted papers

2025

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

ICASSP 2025accepted

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers. The speech quality across the speaker set typically is diverse…

Cited by 0SourceScholar
2023

Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech Synthesis

ICASSP 2023accepted

In recent years, the quality of text-to-speech (TTS) synthesis vastly improved due to deep-learning techniques, with parallel architectures, in particular, providing excellent synthesis quality at fast inference. Training these models usually requires speech recordings, corresponding phoneme-level t…

Cited by 0SourceScholar
2019

Evaluating Salience Representations for Cross-modal Retrieval of Western Classical Music Recordings

ICASSP 2019accepted

In this paper, we consider a cross-modal retrieval scenario of Western classical music. Given a short monophonic musical theme in symbolic notation as query, the objective is to find relevant audio recordings in a database. A major challenge of this retrieval task is the possible difference in the d…

Cited by 0SourceScholar