← Search

Jahn Heymann

9 accepted papers

2023

Towards Accurate and Real-Time End-of-Speech Estimation

ICASSP 2023accepted

We introduce a variant of the endpoint (EP) detection problem in automatic speech recognition (ASR), which we call the end-of-speech (EOS) estimation. Given an utterance, EOS estimation aims to identify the timestamp when the utterance waveform has fully decayed and is then used to measure the EP la…

Cited by 0SourceScholar
2022

Being Greedy Does Not Hurt: Sampling Strategies for End-To-End Speech Recognition

ICASSP 2022accepted

Maximum Likelihood Estimation (MLE) is currently the most common approach to train large scale speech recognition systems. While it has significant practical advantages, MLE exhibits several drawbacks known in literature: training and inference conditions are mismatched and a proxy objective is opti…

Cited by 0SourceScholar
2019

Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASR

ICASSP 2019accepted

Signal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore…

Cited by 0SourceScholar
2018

Performance of Mask Based Statistical Beamforming in a Smart Home Scenario

ICASSP 2018accepted

Mask based statistical beamforming, where signal statistics for the target and the interference gained from masking are used for beamforming, has shown great effectiveness in the two recent CHiME challenges. This idea has sparked interest in the research community and resulted in numerous proposed a…

Cited by 0SourceScholar
2017

Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system

ICASSP 2017accepted

This paper presents an end-to-end training approach for a beamformer-supported multi-channel ASR system. A neural network which estimates masks for a statistically optimum beamformer is jointly trained with a network for acoustic modeling. To update its parameters, we propagate the gradients from th…

Cited by 0SourceScholar
2017

Optimizing neural-network supported acoustic beamforming by algorithmic differentiation

ICASSP 2017accepted

In this paper we show how a neural network for spectral mask estimation for an acoustic beamformer can be optimized by algorithmic differentiation. Using the beamformer output SNR as the objective function to maximize, the gradient is propagated through the beamformer all the way to the neural netwo…

Cited by 0SourceScholar
2016

Neural network based spectral mask estimation for acoustic beamforming

ICASSP 2016accepted

We present a neural network based approach to acoustic beamforming. The network is used to estimate spectral masks from which the Cross-Power Spectral Density matrices of speech and noise are estimated, which in turn are used to compute the beamformer coefficients. The network training is independen…

Cited by 0SourceScholar
2015

Unsupervised adaptation of a denoising autoencoder by Bayesian Feature Enhancement for reverberant asr under mismatch conditions

ICASSP 2015accepted

The parametric Bayesian Feature Enhancement (BFE) and a datadriven Denoising Autoencoder (DA) both bring performance gains in severe single-channel speech recognition conditions. The first can be adjusted to different conditions by an appropriate parameter setting, while the latter needs to be train…

Cited by 0SourceScholar