← Search

Christian Dittmar

4 accepted papers

2025

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

ICASSP 2025accepted

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers. The speech quality across the speaker set typically is diverse…

Cited by 0SourceScholar
2023

Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech Synthesis

ICASSP 2023accepted

In recent years, the quality of text-to-speech (TTS) synthesis vastly improved due to deep-learning techniques, with parallel architectures, in particular, providing excellent synthesis quality at fast inference. Training these models usually requires speech recordings, corresponding phoneme-level t…

Cited by 0SourceScholar
2018

Unifying Local and Global Methods for Harmonic-Percussive Source Separation

ICASSP 2018accepted

This paper addresses the separation of drums from music recordings, a task closely related to harmonic-percussive source separation (HPSS). In previous works, two families of algorithms have been prominently applied to this problem. They are based either on local filtering and diffusion schemes, or…

Cited by 0SourceScholar
2017

Data-driven solo voice enhancement for jazz music retrieval

ICASSP 2017accepted

Retrieving short monophonic queries in music recordings is a challenging research problem in Music Information Retrieval (MIR). In jazz music, given a solo transcription, one retrieval task is to find the corresponding (potentially polyphonic) recording in a music collection. Many conventional syste…

Cited by 0SourceScholar