← Search

Stefan Uhlich

11 accepted papers

2025

Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer

ICASSP 2025accepted

Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired mono…

Cited by 0SourceScholar
2024

SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning

ECCV 2024poster

"Handling distribution shifts from training data, known as out-of-distribution (OOD) generalization, poses a significant challenge in the field of machine learning. While a pre-trained vision-language model like CLIP has demonstrated remarkable zero-shot performance, further adaptation of the model…

2023

Autotts: End-to-End Text-to-Speech Synthesis Through Differentiable Duration Modeling

ICASSP 2023accepted

Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they are not jointly trained. In this paper, we propose a differentiable duration met…

Cited by 0SourceScholar
2023

Improving Self-Supervised Learning for Audio Representations by Feature Diversity and Decorrelation

ICASSP 2023accepted

Self-supervised learning (SSL) has recently shown remarkable results in closing the gap between supervised and unsupervised learning. The idea is to learn robust features that are invariant to distortions of the input data. Despite its success, this idea can suffer from a collapsing issue where the…

Cited by 0SourceScholar
2023

Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects

ICASSP 2023accepted

We propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording.…

Cited by 0SourceScholar
2022

Music Source Separation With Deep Equilibrium Models

ICASSP 2022accepted

While deep neural network-based music source separation (MSS) is very effective and achieves high performance, its model size is often a problem for practical deployment. Deep implicit architectures such as deep equilibrium models (DEQ) were recently proposed, which can achieve higher performance th…

Cited by 0SourceScholar
2021

All For One And One For All: Improving Music Separation By Bridging Networks

ICASSP 2021accepted

This paper proposes several improvements for music separation with deep neural networks (DNNs), namely a multi-domain loss (MDL) and two combination schemes. First, by using MDL we take advantage of the frequency and time domain representation of audio signals. Next, we utilize the relationship amon…

Cited by 0SourceScholar
2020

Mixed Precision DNNs: All you need is a good parametrization

ICLR 2020poster

Efficient deep neural network (DNN) inference on mobile or embedded devices typically involves quantization of the network parameters and activations. In particular, mixed precision networks achieve better performance than networks with homogeneous bitwidth for the same size constraint. Since choosi…

Cited by 0SourcecodeScholar
2017

Improving music source separation based on deep neural networks through data augmentation and network blending

ICASSP 2017accepted

This paper deals with the separation of music into individual instrument tracks which is known to be a challenging problem. We describe two different deep neural network architectures for this task, a feed-forward and a recurrent one, and show that each of them yields themselves state-of-the art res…

Cited by 0SourceScholar
2015

NMF-based blind source separation using a linear predictive coding error clustering criterion

ICASSP 2015accepted

Non-negative matrix factorization (NMF) based sound source separation involves two phases: First, the signal spectrum is decomposed into components which, in a second step, are clustered in order to obtain estimates of the source signal spectra. The major challenge with this approach is the accuracy…

Cited by 0SourceScholar