← Search

Timo Gerkmann

37 accepted papers

2026

ARE MODERN SPEECH ENHANCEMENT SYSTEMS VULNERABLE TO ADVERSARIAL ATTACKS?

ICASSP 2026poster

Machine learning approaches for speech enhancement are becoming increasingly expressive, enabling ever more powerful modifications of input signals. In this paper, we demonstrate that this expressiveness introduces a vulnerability: advanced speech enhancement models can be susceptible to adversarial…

Cited by 0SourcePDFScholar
2025

FlowDec: A flow-based full-band general audio codec with high perceptual quality

ICLR 2025poster

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from…

2025

HRTF Estimation using a Score-based Prior

ICASSP 2025accepted

We present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human speech. The impulse response of the room is estimated along…

Cited by 0SourceScholar
2025

Investigating Training Objectives for Generative Speech Enhancement

ICASSP 2025accepted

Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these framewor…

Cited by 0SourceScholar
2025

Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask Estimation

ICASSP 2025accepted

Due to their robustness and flexibility, neural-driven beamformers are a popular choice for speech separation in challenging environments with a varying amount of simultaneous speakers alongside noise and reverberation. Time-frequency masks and relative directions of the speakers regarding a fixed s…

Cited by 0SourceScholar
2024

A Flexible Online Framework for Projection-Based Stft Phase Retrieval

ICASSP 2024accepted

Several recent contributions in the field of iterative STFT phase retrieval have demonstrated that the performance of the classical Griffin-Lim method can be considerably improved upon. By using the same projection operators as Griffin-Lim, but combining them in innovative ways, these approaches ach…

Cited by 0SourceScholar
2024

EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild Data

ICASSP 2024accepted

Speech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on acted-out datasets and parallel data samples, in this work we…

Cited by 0SourceScholar
2024

Live Iterative Ptychography with Projection-Based Algorithms

ICASSP 2024accepted

In this work, we demonstrate that the ptychographic phase problem can be solved in a live fashion during scanning, while data is still being collected. We propose a generally applicable modification of the widespread projection-based algorithms such as Error Reduction (ER) and Difference Map (DM). T…

Cited by 0SourceScholar
2024

Single and Few-Step Diffusion for Generative Speech Enhancement

ICASSP 2024accepted

Diffusion models have shown promising results in single-channel speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score estimation is called multiple times to solve the i…

Cited by 0SourceScholar
2023

Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech Restoration

ICASSP 2023accepted

Diffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks like speech enhancement and dereverberation. While discriminative models have tra…

Cited by 0SourceScholar
2023

Partially Adaptive Multichannel Joint Reduction of Ego-Noise and Environmental Noise

ICASSP 2023accepted

Human-robot interaction relies on a noise-robust audio processing module capable of estimating target speech from audio recordings impacted by environmental noise, as well as self-induced noise, so-called ego-noise. While external ambient noise sources vary from environment to environment, ego-noise…

Cited by 0SourceScholar
2023

Speech Signal Improvement Using Causal Generative Diffusion Models

ICASSP 2023accepted

In this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processi…

Cited by 0SourceScholar
2022

Customizable End-To-End Optimization Of Online Neural Network-Supported Dereverberation For Hearing Devices

ICASSP 2022accepted

This work focuses on online dereverberation for hearing devices using the weighted prediction error (WPE) algorithm. WPE filtering requires an estimate of the target speech power spectral density (PSD). Recently deep neural networks (DNNs) have been used for this task. However, these approaches opti…

Cited by 0SourceScholar
2022

Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement

ICASSP 2022accepted

Speech enhancement in the time-frequency domain is often performed by estimating a multiplicative mask to extract clean speech. However, most neural network-based methods perform point estimation, i.e., their output consists of a single mask. In this paper, we study the benefits of modeling uncertai…

Cited by 0SourceScholar
2021

Guided Variational Autoencoder for Speech Enhancement with a Supervised Classifier

ICASSP 2021accepted

Recently, variational autoencoders have been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. However, variational autoencoders are trained on clean speech only, which results in a limited ability of extracting the speech signal…

Cited by 0SourceScholar
2021

Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural Network

NeurIPS 2021poster

Recent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired architecture called Fully Recurrent Convolutional Neural Netwo…

2021

Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder

ICASSP 2021accepted

Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to noise presence, especially in low signal-to-noise ratios (SNR…

Cited by 0SourceScholar
2020

Nonlinear Spatial Filtering for Multichannel Speech Enhancement in Inhomogeneous Noise Fields

ICASSP 2020accepted

A common processing pipeline for multichannel speech enhancement is to combine a linear spatial filter with a single-channel postfilter. In fact, it can be shown that such a combination is optimal in the minimum mean square error (MMSE) sense if the noise follows a multivariate Gaussian distribution…

Cited by 0SourceScholar
2020

Robust Robotic Pouring using Audition and Haptics

IROS 2020poster

Robust and accurate estimation of liquid height lies as an essential part of pouring tasks for service robots. However, vision-based methods often fail in occluded conditions while audio-based methods cannot work well in a noisy environment. We instead propose a multimodal pouring network (MP-Net) t…

Cited by 24SourcecodeScholar
2019

An Analysis of Noise-aware Features in Combination with the Size and Diversity of Training Data for DNN-based Speech Enhancement

ICASSP 2019accepted

In this work, the generalization of speech enhancement algorithms based on deep neural networks (DNNs) for training datasets that differ in size and diversity is analyzed. For this, we compare noise aware training (NAT) features and signal-to-noise ratio (SNR) based noise aware training (SNR-NAT) fe…

Cited by 0SourceScholar
2019

Making Sense of Audio Vibration for Liquid Height Estimation in Robotic Pouring

IROS 2019poster

In this paper, we focus on the challenging perception problem in robotic pouring. Most of the existing approaches either leverage visual or haptic information. However, these techniques may suffer from poor generalization performances on opaque containers or concerning measuring precision. To tackle…

Cited by 43SourceScholar
2018

Weighted and Multi-Task Loss for Rare Audio Event Detection

ICASSP 2018accepted

We present in this paper two loss functions tailored for rare audio event detection in audio streams. The weighted loss is designed to tackle the common issue of imbalanced data in background/foreground classification while the multi-task loss enables the networks to simultaneously model the class d…

Cited by 0SourceScholar
2016

BIAS correction methods for adaptive recursive smoothing with applications in noise PSD estimation

ICASSP 2016accepted

Due to the low computational complexity and the low memory consumption, first-order recursive smoothing is a technique often applied to estimate the mean of a random process. For instance, recursive smoothing is used in noise power estimators where adaptively changing smoothing factors are used inst…

Cited by 0SourceScholar
2016

Perceptual and instrumental evaluation of the perceived level of reverberation

ICASSP 2016accepted

Perceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significan…

Cited by 13SourceScholar
2015

Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparison

ICASSP 2015accepted

In this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities…

Cited by 0SourceScholar
2015

Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximation

ICASSP 2015accepted

In many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transf…

Cited by 14SourceScholar
2015

Utilizing spectro-temporal correlations for an improved speech presence probability based noise power estimation

ICASSP 2015accepted

For the enhancement of speech degraded by noise, accurate estimation of the noise power spectral density (PSD) is indispensable, especially if only a single microphone signal is available. Fast and accurate tracking of the noise PSD is particularly challenging in highly non-stationary noise types, s…

Cited by 0SourceScholar