← Search

Jesper Rindom Jensen

29 accepted papers

2026

TOWARDS FAIR ASR FOR SECOND LANGUAGE SPEAKERS USING FAIRNESS PROMPTED FINETUNING

ICASSP 2026poster

In this work, we address the challenge of building fair English ASR systems for second-language speakers. Our analysis of widely used ASR models, Whisper and Seamless-M4T, reveals large fluctuations in word error rate (WER) across 26 accent groups, indicating significant fairness gaps. To mitigate t…

Cited by 0SourcePDFScholar
2025

Advances in Microphone Array Processing and Multichannel Speech Enhancement

ICASSP 2025accepted

This paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas…

Cited by 0SourceScholar
2025

Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking

ICASSP 2025accepted

Performance of sound zone control (SZC) systems deployed in practical scenarios are highly sensitive to the location of the listener(s) and can degrade significantly when listener(s) are moving. This paper presents a robust SZC system that adapts to dynamic changes such as moving listeners and varyi…

Cited by 1SourceScholar
2025

Sound Zone Control Robust To Sound Speed Change

ICASSP 2025accepted

Sound zone control (SZC) implemented using static optimal filters is significantly affected by various perturbations in the acoustic environment, an important one being the fluctuation in the speed of sound, which is in turn influenced by changes in temperature and humidity (TH). This issue arises b…

Cited by 3SourceScholar
2024

Broadband Personal Sound Zone Control in the Presence of Nonlinearities

ICASSP 2024accepted

Existing literature on sound zone control generally consider the signal model to be linear. However, this is seldom true in practice owing to nonlinear distortions arising from the loudspeakers, especially in consumer applications. In this paper, we propose a new signal model for personal sound zone…

Cited by 0SourceScholar
2023

Frequency Bin-Wise Single Channel Speech Presence Probability Estimation Using Multiple DNNS

ICASSP 2023accepted

In this work, we propose a frequency bin-wise method to estimate the single-channel speech presence probability (SPP) with multiple deep neural networks (DNNs) in the short-time Fourier transform domain. Since all frequency bins are typically considered simultaneously as input features for conventio…

Cited by 0SourceScholar
2023

Sparse Bayesian Learning Based Three-Dimensional Imaging for Antenna Array Radar

ICASSP 2023accepted

In recent years, the development of compressed sensing and sparse representation provide us with a broader perspective of three-dimensional (3-D) imaging. In this work, we propose a 3-D imaging method based on a sparse Bayesian learning(SBL) framework for antenna array radar. It solves the problem o…

Cited by 0SourceScholar
2022

Sparse Modeling of The Early Part of Noisy Room Impulse Responses with Sparse Bayesian Learning

ICASSP 2022accepted

A model of a room impulse response (RIR) is useful for a wide range of applications. Typically, the early part of a RIR is sparse, and its sparse structure allows for accurate and simple modeling of the RIR. The existing ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w…

Cited by 0SourceScholar
2019

Quality Control of Voice Recordings in Remote Parkinson's Disease Monitoring Using the Infinite Hidden Markov Model

ICASSP 2019accepted

The performance of voice-based systems for remote monitoring of Parkinson's disease is highly dependent on the degree of adherence of the recordings to the test protocols, which probe for specific symptoms. Identifying segments of the signal that adhere to the protocol assumptions is typically perfo…

Cited by 0SourceScholar
2018

A Parametric Approach for Classification of Distortions in Pathological Voices

ICASSP 2018accepted

In biomedical acoustics, distortion in voice signals, commonly present during acquisition and transmission, adversely affects acoustic features extracted from pathological voice. Information on the type of distortion can help in compensating for its effects. This paper proposes a new approach to det…

Cited by 0SourceScholar
2018

A Supervised Approach to Global Signal-to-Noise Ratio Estimation for Whispered and Pathological Voices

ICASSP 2018accepted

The presence of background noise in signals adversely affects the performance of many speech-based algorithms. Accurate estimation of signal-to-noise-ratio (SNR), as a measure of noise level in a signal, can help in compensating for noise effects. Most existing SNR estimation methods have been devel…

Cited by 0SourceScholar
2018

A Unified Approach to Generating Sound Zones Using Variable Span Linear Filters

ICASSP 2018accepted

Sound zones are typically created using Acoustic Contrast Control (ACC), Pressure Matching (PM), or variations of the two. ACC maximizes the acoustic potential energy contrast between a listening zone and a quiet zone. Although the contrast is maximized, the phase is not controlled. To control both…

Cited by 0SourceScholar
2018

Multipitch Estimation Using Block Sparse Bayesian Learning and Intra-Block Clustering

ICASSP 2018accepted

Pitch estimation is an important task in speech and audio analysis. In this paper, we present a multi-pitch estimation algorithm based on block sparse Bayesian learning and intra-block clustering for speech analysis. A statistical hierarchical model is formulated based on a pitch dictionary with a f…

Cited by 0SourceScholar
2017

Distributed max-SINR speech enhancement with ad hoc microphone arrays

ICASSP 2017accepted

In recent years, signal processing with ad hoc microphone arrays has attracted a lot of attention. Speech enhancement in noisy, interfered, and reverberant environments is one of the problems targeted by ad hoc microphone arrays. Most of the proposed solutions require knowledge of fingerprints, such…

Cited by 0SourceScholar
2017

Estimation of multiple pitches in stereophonic mixtures using a codebook-based approach

ICASSP 2017accepted

In this paper, a method for multi-pitch estimation of stereophonic mixtures of multiple harmonic signals is presented. The method is based on a signal model which takes the amplitude and delay panning parameters of the sources in a stereophonic mixture into account. Furthermore, the method is based…

Cited by 0SourceScholar
2017

Fast harmonic chirp summation

ICASSP 2017accepted

The harmonic chirp signal model has only very recently been introduced for modelling approximately periodic signals with a time-varying fundamental frequency. A number of estimators for the parameters of this model have already been proposed, but they are either inaccurate, non-robust to noise, or v…

Cited by 0SourceScholar
2017

Harmonic minimum mean squared error filters for multichannel speech enhancement

ICASSP 2017accepted

Many state-of-the-art multichannel speech enhancement methods rely on second-order statistics of the desired speech signal, the noise signal, or both. Estimation of those are difficult in practice, resulting in a practical performance that is typically much lower than their potential theoretical per…

Cited by 0SourceScholar
2017

Least 1-norm pole-zero modeling with sparse deconvolution for speech analysis

ICASSP 2017accepted

In this paper, we present a speech analysis method based on sparse pole-zero modeling of speech. Instead of using the all-pole model to approximate the speech production filter, a pole-zero model is used for the combined effect of the vocal tract; radiation at the lips and the glottal pulse shape. M…

Cited by 0SourceScholar
2016

A partitioned approach to signal separation with microphone ad hoc arrays

ICASSP 2016accepted

In this paper, a blind algorithm is proposed for speech enhancement in multi-speaker scenarios, in which interference rejection is the main objective. Here, the ad hoc array is broken into microphone duples which are used to partition the array into local sub-arrays. The core algorithm takes advanta…

Cited by 0SourceScholar
2016

DOA estimation of audio sources in reverberant environments

ICASSP 2016accepted

Reverberation is well-known to have a detrimental impact on many localization methods for audio sources. We address this problem by imposing a model for the early reflections as well as a model for the audio source itself. Using these models, we propose two iterative localization methods that estima…

Cited by 0SourceScholar
2016

Fast and statistically efficient fundamental frequency estimation

ICASSP 2016accepted

Fundamental frequency estimation is a very important task in many applications involving periodic signals. For computational reasons, fast autocorrelation-based estimation methods are often used despite parametric estimation methods having superior estimation accuracy. However, these parametric meth…

Cited by 10SourceScholar
2015

On frequency domain models for TDOA estimation

ICASSP 2015accepted

Time-difference-of-arrival (TDOA) estimation is an important problem in many microphone signal processing applications. Traditionally, this problem is solved by using a cross-correlation method, but in this paper we show that the cross-correlation method is actually a restricted special case of a mu…

Cited by 0SourceScholar
2015

Pitch and TDOA-based localization of acoustic sources with distributed arrays

ICASSP 2015accepted

In this paper, a method for acoustic source localization using distributed microphone arrays based on time-differences of arrival (TDOAs) is presented. The TDOAs are used to estimate the location of an acoustic source using a recently proposed method, based on a 4D parameter space defined by the 3D…

Cited by 1SourceScholar
2015

Pitch estimation and tracking with harmonic emphasis on the acoustic spectrum

ICASSP 2015accepted

In this paper, we use unconstrained frequency estimates (UFEs) from a noisy harmonic signal and propose two methods to estimate and track the pitch over time. We assume that the UFEs are multivariate-normally-distributed random variables, and derive a maximum likelihood (ML) pitch estimator by maxim…

Cited by 0SourceScholar
2015

Pseudo-coherence-based MVDR beamformer for speech enhancement with ad hoc microphone arrays

ICASSP 2015accepted

Speech enhancement with distributed arrays has been met with various methods. On the one hand, data independent methods require information about the position of sensors, so they are not suitable for dynamic geometries. On the other hand, Wiener-based methods cannot assure a distortionless output. T…

Cited by 0SourceScholar