← Search

Reinhold Haeb-Umbach

31 accepted papers

2026

LOOSE COUPLING OF SPECTRAL AND SPATIAL MODELS FOR MULTI-CHANNEL DIARIZATION AND ENHANCEMENT OF MEETINGS IN DYNAMIC ENVIRONMENTS

ICASSP 2026poster

Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. Here, we…

Cited by 0SourcePDFScholar
2025

30+ Years of Source Separation Research: Achievements and Future Challenges

ICASSP 2025accepted

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> anniversary, we review the major contribut…

Cited by 0SourceScholar
2025

Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models

ICASSP 2025accepted

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (vMFMM) for diarization in a joint statistical framework. Through the integr…

Cited by 0SourceScholar
2025

Speech Synthesis along Perceptual Voice Quality Dimensions

ICASSP 2025accepted

While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities (PVQs) recognized by phonetic experts, which are speech properti…

Cited by 0SourceScholar
2024

Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting Scenarios

ICASSP 2024accepted

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions w…

Cited by 0SourceScholar
2023

Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting Diarization

ICASSP 2023accepted

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings…

Cited by 0SourceScholar
2023

On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems

ICASSP 2023accepted

We propose a general framework to compute the word error rate (WER) of ASR systems that process recordings containing multiple speakers at their input and that produce multiple output word sequences (MIMO). Such ASR systems are typically required, e.g., for meeting transcription. We provide an effic…

Cited by 0SourceScholar
2022

On Synchronization of Wireless Acoustic Sensor Networks in the Presence of Time-Varying Sampling Rate Offsets and Speaker Changes

ICASSP 2022accepted

A wireless acoustic sensor network records audio signals with sampling time and sampling rate offsets between the audio streams, if the analog-digital converters (ADCs) of the network devices are not synchronized. Here, we introduce a new sampling rate offset model to simulate time-varying sampling…

Cited by 0SourceScholar
2022

SA-SDR: A Novel Loss Function for Separation of Meeting Style Data

ICASSP 2022accepted

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal perfectly or if the reference signal contains silence, e.g.,…

Cited by 0SourceScholar
2021

Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech Representations

ICASSP 2021accepted

In this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentangl…

Cited by 0SourceScholar
2021

Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation

ICASSP 2021accepted

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine ne…

Cited by 0SourceScholar
2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

ICASSP 2021accepted

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both singlechannel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performan…

Cited by 0SourceScholar
2021

Iterative Geometry Calibration from Distance Estimates for Wireless Acoustic Sensor Networks

ICASSP 2021accepted

In this paper we present an approach to geometry calibration in wireless acoustic sensor networks, whose nodes are assumed to be equipped with a compact microphone array. The proposed approach solely works with estimates of the distances between acoustic sources and the nodes that record these sourc…

Cited by 0SourceScholar
2020

Demystifying TasNet: A Dissecting Approach

ICASSP 2020accepted

In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the gains of the time-domain audio separation network (TasNet) approach by gradually replacing components of an utterance-leve…

Cited by 0SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2020

Jointly Optimal Dereverberation and Beamforming

ICASSP 2020accepted

We previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a Weighted Prediction Error (WPE) dereverberation filter and a conventional Minimum-Pow…

Cited by 0SourceScholar
2019

All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis

ICASSP 2019accepted

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant…

Cited by 0SourceScholar
2019

Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASR

ICASSP 2019accepted

Signal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore…

Cited by 0SourceScholar
2019

Unsupervised Training of a Deep Clustering Model for Multichannel Blind Source Separation

ICASSP 2019accepted

We propose a training scheme to train neural network-based source separation algorithms from scratch when parallel clean data is unavailable. In particular, we demonstrate that an unsupervised spatial clustering algorithm is sufficient to guide the training of a deep clustering system. We argue that…

Cited by 0SourceScholar
2018

Deep Attractor Networks for Speaker Re-Identification and Blind Source Separation

ICASSP 2018accepted

Deep clustering (DC) and deep attractor networks (DANs) are a data-driven way to monaural blind source separation. Both approaches provide astonishing single channel performance but have not yet been generalized to block-online processing. When separating speech in a continuous stream with a block-o…

Cited by 0SourceScholar
2018

Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source Separation

ICASSP 2018accepted

Deep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on bloc…

Cited by 0SourceScholar
2018

Exploring Practical Aspects of Neural Mask-Based Beamforming for Far-Field Speech Recognition

ICASSP 2018accepted

This work examines acoustic beamformers employing neural networks (NNs) for mask prediction as front -end for automatic speech recognition (ASR) systems for practical scenarios like voice-enabled home devices. To test the versatility of the mask predicting network, the system is evaluated with diffe…

Cited by 0SourceScholar
2017

A generalized log-spectral amplitude estimator for single-channel speech enhancement

ICASSP 2017accepted

The benefits of both a logarithmic spectral amplitude (LSA) estimation and a modeling in a generalized spectral domain (where short-time amplitudes are raised to a generalized power exponent, not restricted to magnitude or power spectrum) are combined in this contribution to achieve a better tradeof…

Cited by 0SourceScholar
2017

Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system

ICASSP 2017accepted

This paper presents an end-to-end training approach for a beamformer-supported multi-channel ASR system. A neural network which estimates masks for a statistically optimum beamformer is jointly trained with a network for acoustic modeling. To update its parameters, we propagate the gradients from th…

Cited by 0SourceScholar
2017

Optimizing neural-network supported acoustic beamforming by algorithmic differentiation

ICASSP 2017accepted

In this paper we show how a neural network for spectral mask estimation for an acoustic beamformer can be optimized by algorithmic differentiation. Using the beamformer output SNR as the objective function to maximize, the gradient is propagated through the beamformer all the way to the neural netwo…

Cited by 0SourceScholar
2016

Blind speech separation based on complex spherical k-mode clustering

ICASSP 2016accepted

We present an algorithm for clustering complex-valued unit length vectors on the unit hypersphere, which we call complex spherical k-mode clustering, as it can be viewed as a generalization of the spherical k-means algorithm to normalized complex-valued vectors. We show how the proposed algorithm ca…

Cited by 0SourceScholar
2016

Neural network based spectral mask estimation for acoustic beamforming

ICASSP 2016accepted

We present a neural network based approach to acoustic beamforming. The network is used to estimate spectral masks from which the Cross-Power Spectral Density matrices of speech and noise are estimated, which in turn are used to compute the beamformer coefficients. The network training is independen…

Cited by 0SourceScholar
2015

Aligning training modelswith smartphone properties in WiFi fingerprinting based indoor localization

ICASSP 2015accepted

We are concerned with the so-called fingerprinting method for WiFi-based indoor positioning, where the measured received signal strength index (RSSI) is compared with training data to come up with an estimate of the user's location. We introduce a method for adapting the trained models to the statis…

Cited by 0SourceScholar
2015

Source counting in speech mixtures by nonparametric Bayesian estimation of an infinite Gaussian mixture model

ICASSP 2015accepted

In this paper we present a source counting algorithm to determine the number of speakers in a speech mixture. In our proposed method, we model the histogram of estimated directions of arrival with a non-parametric Bayesian infinite Gaussian mixture model. As an alternative to classical model selecti…

Cited by 0SourceScholar
2015

Unsupervised adaptation of a denoising autoencoder by Bayesian Feature Enhancement for reverberant asr under mismatch conditions

ICASSP 2015accepted

The parametric Bayesian Feature Enhancement (BFE) and a datadriven Denoising Autoencoder (DA) both bring performance gains in severe single-channel speech recognition conditions. The first can be adjusted to different conditions by an appropriate parameter setting, while the latter needs to be train…

Cited by 0SourceScholar