← Search

Thilo von Neumann

5 accepted papers

2023

On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems

ICASSP 2023accepted

We propose a general framework to compute the word error rate (WER) of ASR systems that process recordings containing multiple speakers at their input and that produce multiple output word sequences (MIMO). Such ASR systems are typically required, e.g., for meeting transcription. We provide an effic…

Cited by 0SourceScholar
2022

SA-SDR: A Novel Loss Function for Separation of Meeting Style Data

ICASSP 2022accepted

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal perfectly or if the reference signal contains silence, e.g.,…

Cited by 0SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2019

All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis

ICASSP 2019accepted

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant…

Cited by 0SourceScholar
2018

Deep Attractor Networks for Speaker Re-Identification and Blind Source Separation

ICASSP 2018accepted

Deep clustering (DC) and deep attractor networks (DANs) are a data-driven way to monaural blind source separation. Both approaches provide astonishing single channel performance but have not yet been generalized to block-online processing. When separating speech in a continuous stream with a block-o…

Cited by 0SourceScholar