← Search

Jesper Jensen

30 accepted papers

2025

Deep Feedback Cancellation for Hearing Aids with Improved System Stability and Sound Quality

ICASSP 2025accepted

Acoustic feedback cancellation is an important task in audio processing systems, aiming to mitigate the effects of feedback loops on system stability and sound quality. State-of-the-art methods rely on adaptive filtering algorithms and face challenges in balancing between rapid convergence and low s…

Cited by 0SourceScholar
2024

Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks

ICASSP 2024accepted

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an enco…

Cited by 0SourceScholar
2024

Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler

ICASSP 2024accepted

Diffusion models are a new class of generative models that have recently been applied to speech enhancement successfully. Previous works have demonstrated their superior performance in mismatched conditions compared to state-of-the art discriminative models. However, this was investigated with a sin…

Cited by 0SourceScholar
2024

Self-Supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions

ICASSP 2024accepted

In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory (LSTM)-encoder using the autoregressive predictive coding (APC…

Cited by 0SourceScholar
2024

Speaker Adaptation For Enhancement Of Bone-Conducted Speech

ICASSP 2024accepted

Deep neural network (DNN)-based speech enhancement models often face challenges in maintaining their performance for speakers not encountered during training. This challenge is exacerbated in applications such as enhancement and bandwidth extension of bone-conducted speech, where the distortion exhi…

Cited by 0SourceScholar
2024

Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal

ICASSP 2024accepted

Speech enhancement in hearing aids (HAs) can take advantage of a wireless remote microphone (RM) having a better signal-to-noise ratio than the HA microphones. However, using the RM effectively is complicated by the time delay between the acoustic and wireless signals. Methods in the literature assu…

Cited by 0SourceScholar
2023

Distributed Adaptive Norm Estimation for Blind System Identification in Wireless Sensor Networks

ICASSP 2023accepted

Distributed signal-processing algorithms in (wireless) sensor networks often aim to decentralize processing tasks to reduce communication cost and computational complexity or avoid reliance on a single device (i.e., fusion center) for processing. In this contribution, we extend a distributed adaptiv…

Cited by 0SourceScholar
2023

Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting

ICASSP 2023accepted

In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels i…

Cited by 0SourceScholar
2022

Joint Far- and Near-End Speech Intelligibility Enhancement Based on the Approximated Speech Intelligibility Index

ICASSP 2022accepted

This paper considers speech enhancement of signals picked up in one noisy environment which must be presented to a listener in another noisy environment. Recently, it has been shown that an optimal solution to this problem requires the consideration of the noise sources in both environments jointly.…

Cited by 0SourceScholar
2021

Audio-Visual Speech Inpainting with Deep Learning

ICASSP 2021accepted

In this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliable audio context and uncorrupted visual information. Recent work focuses solely on audio-only methods and generally aims…

Cited by 31SourceScholar
2021

Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic Beamforming

ICASSP 2021accepted

Acoustic beamforming is crucial for many applications where ex-traction of a target signal from a noisy environment is required. In order to implement practical beamformers, e.g. the multichannel Wiener filter (MWF), estimation of the target and noise power spectral densities (PSDs), and the relativ…

Cited by 8SourceScholar
2020

A Constrained Maximum Likelihood Estimator of Speech and Noise Spectra with Application to Multi-Microphone Noise Reduction

ICASSP 2020accepted

One of the challenges with the implementation of multi-microphone noise reduction systems in practical applications lies in the need for the knowledge of the speech and noise covariance matrices. Recently, a method based on Maximum Likelihood (ML) estimation addressed this problem. Despite its relat…

Cited by 0SourceScholar
2020

A Neural Network for Monaural Intrusive Speech Intelligibility Prediction

ICASSP 2020accepted

Monaural intrusive speech intelligibility prediction (SIP) methods aim to predict the speech intelligibility (SI) of a single-microphone noisy and/or processed speech signal using the underlying clean speech signal. In the present work, we propose a neural network for monaural intrusive SIP. The pro…

Cited by 0SourceScholar
2020

Maximum Likelihood Estimation of the Interference-Plus-Noise Cross Power Spectral Density Matrix for Own Voice Retrieval

ICASSP 2020accepted

In headset and hearing aid applications, it is of interest to retrieve the user's own voice in a noisy environment, e.g. for telephony applications. To do so, the cross-power spectral density (CPSD) of the interference-plus-noise is required. In this paper, a novel maximum likelihood (ML) estimator…

Cited by 0SourceScholar
2019

A Novel Binaural Beamforming Scheme with Low Complexity Minimizing Binaural-cue Distortions

ICASSP 2019accepted

While the majority of binaural beamformers aim to minimize the output noise power while (approximately) preserving the binaural cues of the sources using constraints, we propose in this paper to minimize the binaural-cue distortions of the sources in the acoustic scene, such that the output noise po…

Cited by 0SourceScholar
2019

Effects of Lombard Reflex on the Performance of Deep-learning-based Audio-visual Speech Enhancement Systems

ICASSP 2019accepted

Humans tend to change their way of speaking when they are immersed in a noisy environment, a reflex known as Lombard effect. Current speech enhancement systems based on deep learning do not usually take into account this change in the speaking style, because they are trained with neutral (non-Lombar…

Cited by 0SourceScholar
2019

On Training Targets and Objective Functions for Deep-learning-based Audio-visual Speech Enhancement

ICASSP 2019accepted

Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the AV-SE task in a supervised manner. In this context, the choic…

Cited by 0SourceScholar
2018

Monaural Speech Enhancement Using Deep Neural Networks by Maximizing a Short-Time Objective Intelligibility Measure

ICASSP 2018accepted

In this paper we propose a Deep Neural Network (D NN) based Speech Enhancement (SE) system that is designed to maximize an approximation of the Short-Time Objective Intelligibility (STOI) measure. We formalize an approximate-STOI cost function and derive analytical expressions for the gradients requ…

Cited by 0SourceScholar
2017

A non-intrusive Short-Time Objective Intelligibility measure

ICASSP 2017accepted

We propose a non-intrusive intelligibility measure for noisy and non-linearly processed speech, i.e. a measure which can predict intelligibility from a degraded speech signal without requiring a clean reference signal. The proposed measure is based on the Short-Time Objective Intelligibility (STOI)…

Cited by 0SourceScholar
2017

Permutation invariant training of deep models for speaker-independent multi-talker speech separation

ICASSP 2017accepted

We propose a novel deep learning training criterion, named permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from the multi-class regression technique and the deep clustering (DPCL) technique, our nov…

Cited by 0SourceScholar
2016

A method for predicting the intelligibility of noisy and non-linearly enhanced binaural speech

ICASSP 2016accepted

We propose and evaluate a binaural speech intelligibility measure. The measure is a binaural extension of the Short-Time Objective Intelligibility (STOI) measure and focuses on predicting the intelligibility of noisy speech which has been enhanced by a speech processing algorithm (e.g. in a hearing…

Cited by 0SourceScholar
2016

Improved multi-microphone noise reduction preserving binaural cues

ICASSP 2016accepted

We propose a new multi-microphone noise reduction technique for binaural cue preservation of the desired source and the interferers. This method is based on the linearly constrained minimum variance (LCMV) framework, where the constraints are used for the binaural cue preservation of the desired sou…

Cited by 0SourceScholar
2016

Informed Direction of Arrival estimation using a spherical-head model for Hearing Aid applications

ICASSP 2016accepted

In this paper, we propose a Direction of Arrival (DoA) estimator for a Hearing Aid System (HAS) which can connect to a wireless microphone worn by a target talker. The wireless microphone "informs" the HAS about the almost noise-free content of the target sound, and the proposed DoA estimator uses t…

Cited by 7SourceScholar
2016

Maximum likelihood PSD estimation for speech enhancement in reverberant and noisy conditions

ICASSP 2016accepted

We propose a novel Power Spectral Density (PSD) estimator for multi-microphone systems operating in reverberant and noisy conditions. The estimator is derived using the maximum likelihood approach and is based on a blocked and pre-whitened additive signal model. The intended application of the estim…

Cited by 0SourceScholar
2015

A simple modification to facilitate robust generalized sidelobe canceller for hearing aids

ICASSP 2015accepted

This work focuses on an adaptive beamformer in a hearing aid application using a generalized sidelobe canceller structure (GSC). In this application, the constraint and blocking matrices in the GSC structure are specifically designed using an estimate of the transfer functions between the target sou…

Cited by 0SourceScholar
2015

Analysis of beamformer directed single-channel noise reduction system for hearing aid applications

ICASSP 2015accepted

We study multi-microphone noise reduction systems consisting of a beamformer and a single-channel (SC) noise reduction stage. In particular, we present and analyse a maximum likelihood (ML) method for jointly estimating the target and noise power spectral densities (psd's) entering the SC filter. We…

Cited by 33SourceScholar
2015

Maximum likelihood approach to "informed" Sound Source Localization for Hearing Aid applications

ICASSP 2015accepted

Most state-of-the-art Sound Source Localization (SSL) algorithms have been proposed for applications which are “uninformed” about the target sound content; however, utilizing a wireless microphone worn by a target talker, enables recent Hearing Aid Systems (HASs) to access to an almost noise-free so…

Cited by 0SourceScholar
2015

Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparison

ICASSP 2015accepted

In this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities…

Cited by 0SourceScholar
2015

On the influence of microphone array geometry on HRTF-based Sound Source Localization

ICASSP 2015accepted

The direction dependence of Head Related Transfer Functions (HRTFs) forms the basis for HRTF-based Sound Source Localization (SSL) algorithms. In this paper, we show how spectral similarities of the HRTFs of different directions in the horizontal plane influence performance of HRTF-based SSL algorit…

Cited by 0SourceScholar
2015

Speech reinforcement in noisy reverberant conditions under an approximation of the short-time SII

ICASSP 2015accepted

While most contributions on speech reinforcement only consider the presence of environmental noise, late reverberation can also severely degrade the intelligibility of speech. In this paper we address the problem of speech reinforcement in noisy and reverberant environments. We use a short-time vers…

Cited by 0SourceScholar