← Search

Lukas Drude

13 accepted papers

2024

Promptformer: Prompted Conformer Transducer for ASR

ICASSP 2024accepted

Context cues carry information which can improve multi-turn interactions in automatic speech recognition (ASR) systems. In this paper, we introduce a novel mechanism inspired by hyper-prompting to fuse textual context with acoustic representations in the attention mechanism. Results on a test set wi…

Cited by 0SourceScholar
2020

Demystifying TasNet: A Dissecting Approach

ICASSP 2020accepted

In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the gains of the time-domain audio separation network (TasNet) approach by gradually replacing components of an utterance-leve…

Cited by 66SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2019

Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASR

ICASSP 2019accepted

Signal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore…

Cited by 0SourceScholar
2019

Unsupervised Training of a Deep Clustering Model for Multichannel Blind Source Separation

ICASSP 2019accepted

We propose a training scheme to train neural network-based source separation algorithms from scratch when parallel clean data is unavailable. In particular, we demonstrate that an unsupervised spatial clustering algorithm is sufficient to guide the training of a deep clustering system. We argue that…

Cited by 0SourceScholar
2018

Deep Attractor Networks for Speaker Re-Identification and Blind Source Separation

ICASSP 2018accepted

Deep clustering (DC) and deep attractor networks (DANs) are a data-driven way to monaural blind source separation. Both approaches provide astonishing single channel performance but have not yet been generalized to block-online processing. When separating speech in a continuous stream with a block-o…

Cited by 0SourceScholar
2018

Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source Separation

ICASSP 2018accepted

Deep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on bloc…

Cited by 4SourceScholar
2018

Listening to Each Speaker One by One with Recurrent Selective Hearing Networks

ICASSP 2018accepted

Deep learning-based single-channel source separation algorithms are currently being actively investigated. Among them, Deep Clustering (DC) and Deep Attractor Networks (DANs) have made it possible to separate an arbitrary number of speakers. In particular, they cleverly combine a neural network and…

Cited by 0SourceScholar
2017

Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system

ICASSP 2017accepted

This paper presents an end-to-end training approach for a beamformer-supported multi-channel ASR system. A neural network which estimates masks for a statistically optimum beamformer is jointly trained with a network for acoustic modeling. To update its parameters, we propagate the gradients from th…

Cited by 0SourceScholar
2017

Optimizing neural-network supported acoustic beamforming by algorithmic differentiation

ICASSP 2017accepted

In this paper we show how a neural network for spectral mask estimation for an acoustic beamformer can be optimized by algorithmic differentiation. Using the beamformer output SNR as the objective function to maximize, the gradient is propagated through the beamformer all the way to the neural netwo…

Cited by 0SourceScholar
2016

Blind speech separation based on complex spherical k-mode clustering

ICASSP 2016accepted

We present an algorithm for clustering complex-valued unit length vectors on the unit hypersphere, which we call complex spherical k-mode clustering, as it can be viewed as a generalization of the spherical k-means algorithm to normalized complex-valued vectors. We show how the proposed algorithm ca…

Cited by 0SourceScholar
2016

Neural network based spectral mask estimation for acoustic beamforming

ICASSP 2016accepted

We present a neural network based approach to acoustic beamforming. The network is used to estimate spectral masks from which the Cross-Power Spectral Density matrices of speech and noise are estimated, which in turn are used to compute the beamformer coefficients. The network training is independen…

Cited by 0SourceScholar
2015

Source counting in speech mixtures by nonparametric Bayesian estimation of an infinite Gaussian mixture model

ICASSP 2015accepted

In this paper we present a source counting algorithm to determine the number of speakers in a speech mixture. In our proposed method, we model the histogram of estimated directions of arrival with a non-parametric Bayesian infinite Gaussian mixture model. As an alternative to classical model selecti…

Cited by 0SourceScholar