← Search

Christoph Böddeker

14 accepted papers

2024

Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting Scenarios

ICASSP 2024accepted

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions w…

Cited by 0SourceScholar
2023

Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting Diarization

ICASSP 2023accepted

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings…

Cited by 0SourceScholar
2023

On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems

ICASSP 2023accepted

We propose a general framework to compute the word error rate (WER) of ASR systems that process recordings containing multiple speakers at their input and that produce multiple output word sequences (MIMO). Such ASR systems are typically required, e.g., for meeting transcription. We provide an effic…

Cited by 0SourceScholar
2023

Reverberation as Supervision For Speech Separation

ICASSP 2023accepted

This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separation required the synthesis of mixtures of mixtures or assumed the existence of a teacher model, making them difficult to…

Cited by 0SourceScholar
2022

SA-SDR: A Novel Loss Function for Separation of Meeting Style Data

ICASSP 2022accepted

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal perfectly or if the reference signal contains silence, e.g.,…

Cited by 0SourceScholar
2021

Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation

ICASSP 2021accepted

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine ne…

Cited by 0SourceScholar
2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

ICASSP 2021accepted

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both singlechannel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performan…

Cited by 0SourceScholar
2020

Demystifying TasNet: A Dissecting Approach

ICASSP 2020accepted

In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the gains of the time-domain audio separation network (TasNet) approach by gradually replacing components of an utterance-leve…

Cited by 0SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2020

Jointly Optimal Dereverberation and Beamforming

ICASSP 2020accepted

We previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a Weighted Prediction Error (WPE) dereverberation filter and a conventional Minimum-Pow…

Cited by 0SourceScholar
2018

Exploring Practical Aspects of Neural Mask-Based Beamforming for Far-Field Speech Recognition

ICASSP 2018accepted

This work examines acoustic beamformers employing neural networks (NNs) for mask prediction as front -end for automatic speech recognition (ASR) systems for practical scenarios like voice-enabled home devices. To test the versatility of the mask predicting network, the system is evaluated with diffe…

Cited by 0SourceScholar
2017

Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system

ICASSP 2017accepted

This paper presents an end-to-end training approach for a beamformer-supported multi-channel ASR system. A neural network which estimates masks for a statistically optimum beamformer is jointly trained with a network for acoustic modeling. To update its parameters, we propagate the gradients from th…

Cited by 0SourceScholar
2017

Optimizing neural-network supported acoustic beamforming by algorithmic differentiation

ICASSP 2017accepted

In this paper we show how a neural network for spectral mask estimation for an acoustic beamformer can be optimized by algorithmic differentiation. Using the beamformer output SNR as the objective function to maximize, the gradient is propagated through the beamformer all the way to the neural netwo…

Cited by 0SourceScholar
2016

Blind speech separation based on complex spherical k-mode clustering

ICASSP 2016accepted

We present an algorithm for clustering complex-valued unit length vectors on the unit hypersphere, which we call complex spherical k-mode clustering, as it can be viewed as a generalization of the spherical k-means algorithm to normalized complex-valued vectors. We show how the proposed algorithm ca…

Cited by 0SourceScholar