← Search

Simon Doclo

46 accepted papers

2026

REFERENCE MICROPHONE SELECTION FOR GUIDED SOURCE SEPARATION BASED ON THE NORMALIZED L-P NORM

ICASSP 2026poster

Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may have a large influence on the quality of the output signal…

Cited by 0SourcePDFScholar
2025

Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers

ICASSP 2025accepted

To estimate the direction of arrival (DOA) of multiple speakers, subspace-based prototype transfer function matching methods such as multiple signal classification (MUSIC) or relative transfer function (RTF) vector matching are commonly employed. In general, these methods require calibrated micropho…

Cited by 0SourceScholar
2025

Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone

ICASSP 2025accepted

Hearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user’s own voice in a noisy environment. In this scenario, own voice reconstruction (OVR) is essential for enhancing the quality…

Cited by 0SourceScholar
2024

Active Learning for Sound Event Classification Using Bayesian Neural Networks with Gaussian Variational Posterior

ICASSP 2024accepted

Manual annotation of audio material is cumbersome. Active learning aims at minimizing the annotation effort by iteratively selecting an acquisition batch of unlabeled data, asking a human to annotate the selected data and re-training a classifier until an annotation budget is depleted. In this paper…

Cited by 3SourceScholar
2024

Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks

ICASSP 2024accepted

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an enco…

Cited by 0SourceScholar
2024

Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple Speakers

ICASSP 2024accepted

To estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different frequencies can be combined. According to how the SPS are combi…

Cited by 0SourceScholar
2024

Effect of Target Signals and Delays on Spatially Selective Active Noise Control for Open-Fitting Hearables

ICASSP 2024accepted

Spatially selective active noise control (ANC) hearables are designed to reduce unwanted noise from certain directions while preserving desired sounds from other directions. In previous studies, the target signal has been defined either as the delayed desired component in one of the reference microp…

Cited by 0SourceScholar
2024

Microphone Subset Selection for the Weighted Prediction Error Algorithm Using a Group Sparsity Penalty

ICASSP 2024accepted

Reverberation can severely degrade the quality of speech signals recorded using microphones in an enclosure. In acoustic sensor networks with spatially distributed microphones, a similar dereverberation performance may be achieved using only a subset of all available microphones. Using the popular c…

Cited by 0SourceScholar
2024

Multi-Microphone Noise Data Augmentation for DNN-Based Own Voice Reconstruction for Hearables in Noisy Environments

ICASSP 2024accepted

Hearables with integrated microphones may offer communication benefits in noisy working environments, e.g. by transmitting the recorded own voice of the user. Systems aiming at reconstructing the clean and full-bandwidth own voice from noisy microphone recordings are often based on supervised learni…

Cited by 10SourceScholar
2023

Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting A Calibrated External Microphone Array

ICASSP 2023accepted

Recently, a relative transfer function (RTF) vector-based method has been proposed to estimate the direction of arrival (DOA) of a target speaker for a binaural hearing aid setup, assuming the availability of external microphones. This method exploits the external microphones to estimate the RTF vec…

Cited by 5SourceScholar
2023

Dereverberation in Acoustic Sensor Networks Using weighted Prediction Error with Microphone-Dependent Prediction Delays

ICASSP 2023accepted

In the last decades several multi-microphone speech dereverberation algorithms have been proposed, among which the weighted prediction error (WPE) algorithm. In the WPE algorithm, a prediction delay is required to reduce the correlation between the prediction signals and the direct component in the…

Cited by 5SourceScholar
2023

Geometry-Aware DOA Estimation Using a Deep Neural Network with Mixed-Data Input Features

ICASSP 2023accepted

Unlike model-based direction of arrival (DoA) estimation algorithms, supervised learning-based DoA estimation algorithms based on deep neural networks (DNNs) are usually trained for one specific microphone array geometry, resulting in poor performance when applied to a different array geometry. In t…

Cited by 0SourceScholar
2022

Optimization of a Fixed Virtual Sensing Feedback ANC Controller For In-Ear Headphones with Multiple Loudspeakers

ICASSP 2022accepted

In this paper we consider an in-ear headphone equipped with an inner microphone and multiple loudspeakers and we propose an optimization procedure with a convex objective function to derive a fixed multi-loudspeaker ANC controller aiming at minimizing the sound pressure at the ear drum. Based on the…

Cited by 0SourceScholar
2020

DNN-Based Speech Presence Probability Estimation for Multi-Frame Single-Microphone Speech Enhancement

ICASSP 2020accepted

Multi-frame approaches for single-microphone speech enhancement, e.g., the multi-frame minimum-power-distortionless-response (MFMPDR) filter, are able to exploit speech correlations across neighboring time frames. In contrast to single-frame approaches such as the Wiener gain, it has been shown that…

Cited by 0SourceScholar
2020

Exploiting Periodicity Features for Joint Detection and DOA Estimation of Speech Sources Using Convolutional Neural Networks

ICASSP 2020accepted

While many algorithms deal with direction of arrival (DOA) estimation and voice activity detection (VAD) as two separate tasks, only a small number of data-driven methods have addressed these two tasks jointly. In this paper, a multi-input single-output convolutional neural network (CNN) is proposed…

Cited by 0SourceScholar
2020

Improving Auditory Attention Decoding Performance of Linear and Non-Linear Methods using State-Space Model

ICASSP 2020accepted

Identifying the target speaker in hearing aid applications is crucial to improve speech understanding. Recent advances in electroencephalography (EEG) have shown that it is possible to identify the target speaker from single-trial EEG recordings using auditory attention decoding (AAD) methods. AAD m…

Cited by 0SourceScholar
2020

Subspace-Based Speech Correlation Vector Estimation for Single-Microphone Multi-Frame MVDR Filtering

ICASSP 2020accepted

Aiming at exploiting the speech correlation across consecutive timeframes in the short-time Fourier transform domain, the multi-frame minimum variance distortionless response (MFMVDR) filter for single-microphone speech enhancement has been proposed. This filter is designed to avoid speech distortio…

Cited by 13SourceScholar
2019

Joint Estimation of RETF Vector and Power Spectral Densities for Speech Enhancement Based on Alternating Least Squares

ICASSP 2019accepted

The multi-channel Wiener filter (MWF) is a well-known multi-microphone speech enhancement technique, aiming at improving the quality of the recorded speech signals in noisy and reverberant environments. Assuming that reverberation and ambient noise can be modeled as a diffuse sound field and the spa…

Cited by 0SourceScholar
2019

RTF-steered Binaural MVDR Beamforming Incorporating an External Microphone for Dynamic Acoustic Scenarios

ICASSP 2019accepted

A well-known binaural noise reduction algorithm is the binaural minimum variance distortionless response beamformer, which can be steered using the relative transfer function (RTF) vectors of the desired source. In this paper, we consider the recently proposed spatial coherence (SC) method to estima…

Cited by 0SourceScholar
2018

Complexity Reduction of Eigenvalue Decomposition-Based Diffuse Power Spectral Density Estimators Using the Power Method

ICASSP 2018accepted

In noisy and reverberant environments speech enhancement techniques such as the multi-channel Wiener filter (MWF) can be used to improve speech quality and intelligibility. Assuming that reverberation and ambient noise can be modeled as diffuse sound fields, such techniques require an estimate of th…

Cited by 0SourceScholar
2018

EEG-Based Auditory Attention Decoding Using Steerable Binaural Superdirective Beamformer

ICASSP 2018accepted

During the last decades significant progress in multi-microphone speech enhancement algorithms has been made for hearing aids. However, the performance of many algorithms depends on identifying the target speaker to be enhanced. To identify the target speaker from single-trial EEG recordings in an a…

Cited by 0SourceScholar
2018

Joint Late Reverberation and Noise Power Spectral Density Estimation in a Spatially Homogeneous Noise Field

ICASSP 2018accepted

Many multi-channel dereverberation and noise reduction techniques such as the multi-channel Wiener filter (MWF) require an estimate of the late reverberation and noise power spectral densities (PSDs). State-of-the-art multi-channel methods for estimating the late reverberation PSD typically assume t…

Cited by 0SourceScholar
2017

Comparison of two binaural beamforming approaches for hearing aids

ICASSP 2017accepted

Beamforming algorithms in binaural hearing aids are crucial to improve speech understanding in background noise for hearing impaired persons. In this study, we compare and evaluate the performance of two recently proposed minimum variance (MV) beamforming approaches for binaural hearing aids. The bi…

Cited by 0SourceScholar
2017

Measuring, modelling and predicting perceived reverberation

ICASSP 2017accepted

This paper investigates the relationship between the perceived level of reverberation and parameters measured from the room impulse response (RIR), as well as the design of an instrumental measure that predicts this perceived level. We first present the results of an experimental listening test cond…

Cited by 0SourceScholar
2017

Null-steering beamformer for acoustic feedback cancellation in a multi-microphone earpiece optimizing the maximum stable gain

ICASSP 2017accepted

Commonly adaptive filters are used to reduce the acoustic feedback in hearing aids. While theoretically allowing for perfect cancellation of the feedback signal, in practice the adaptive filter solution is typically biased due to the closed-loop hearing aid system. In contrast to conventional behind…

Cited by 0SourceScholar
2017

Proportionate NLMS for adaptive feedback control in hearing aids

ICASSP 2017accepted

The proportionate normalized least-mean-squares (PNLMS) algorithm is commonly used in acoustic echo cancellation (AEC) context. It provides faster initial convergence and tracking rates compared to the NLMS algorithm for the case of sparse echo impulse responses. The improved PNLMS algorithm (IPNLMS…

Cited by 0SourceScholar
2016

Auditory attention decoding with EEG recordings using noisy acoustic reference signals

ICASSP 2016accepted

To decode auditory attention from electroencephalography (EEG) recordings in a cocktail-party scenario with two competing speakers a least-squares method has recently been proposed, showing a promising decoding accuracy. This method however requires the clean speech signals of both the attended and…

Cited by 0SourceScholar
2016

Extensions of the binaural MWF with interference reduction preserving the binaural cues of the interfering source

ICASSP 2016accepted

Recently, an extension of the binaural multichannel Wiener filter (BMWF), referred to as BMWF-IRo, was presented in which an interference rejection constraint was added to the BMWF cost function. Although the BMWF-IRo aims to entirely suppress the interfering source, residual interfering sources (as…

Cited by 0SourceScholar
2016

Improving adaptive feedback cancellation in hearing aids using an affine combination of filters

ICASSP 2016accepted

In adaptive feedback cancellation an adaptive filter is used to model the acoustic feedback path between the hearing aid loudspeaker and the microphone. An important parameter for adaptive filters is the step-size, providing a trade-off between fast convergence and low steady-state misalignment. In…

Cited by 0SourceScholar
2016

Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aids

ICASSP 2016accepted

Besides noise reduction, an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. For the desired speech source and an interfering source, e.g., competing speaker, this can be achieved by preserving their relative transfer functi…

Cited by 0SourceScholar
2016

Maximum likelihood PSD estimation for speech enhancement in reverberant and noisy conditions

ICASSP 2016accepted

We propose a novel Power Spectral Density (PSD) estimator for multi-microphone systems operating in reverberant and noisy conditions. The estimator is derived using the maximum likelihood approach and is based on a blocked and pre-whitened additive signal model. The intended application of the estim…

Cited by 0SourceScholar
2016

Perceptual and instrumental evaluation of the perceived level of reverberation

ICASSP 2016accepted

Perceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significan…

Cited by 13SourceScholar
2016

Robust sparsity-promoting acoustic multi-channel equalization for speech dereverberation

ICASSP 2016accepted

This paper presents a novel signal-dependent method to increase the robustness of acoustic multi-channel equalization techniques against room impulse response (RIR) estimation errors. Aiming at obtaining an output signal which better resembles a clean speech signal, we propose to extend the acoustic…

Cited by 0SourceScholar
2015

Binaural multichannel Wiener filter with directional interference rejection

ICASSP 2015accepted

In this paper we consider an acoustic scenario with a desired source and a directional interference picked up by hearing devices in a noisy and reverberant environment. We present an extension of the binaural multichannel Wiener filter (BMWF), by adding an interference rejection constraint to its co…

Cited by 0SourceScholar
2015

Common part estimation of acoustic feedback paths in hearing aids optimizing maximum stable gain

ICASSP 2015accepted

The computational complexity and convergence speed of adaptive feedback cancellation algorithms depend on the number of adaptive parameters used to model the acoustic feedback path. To reduce the number of adaptive parameters it has been proposed to decompose the acoustic feedback path as the convol…

Cited by 0SourceScholar
2015

Curvature-based optimization of the trade-off parameter in the speech distortion weighted multichannel wiener filter

ICASSP 2015accepted

The objective of the speech distortion weighted multichannel Wiener filter (MWF) is to reduce background noise while controlling speech distortion. This can be achieved by means of a trade-off parameter, hence, selecting an optimal trade-off parameter is of crucial importance. Aiming at incorporatin…

Cited by 0SourceScholar
2015

Individualizing a monaural beamformer for cochlear implant users

ICASSP 2015accepted

Speech intelligibility in noisy environments is still quite limited for cochlear implant (CI) users. Classical beamformers such as the Generalized Sidelobe Canceller (GSC) can provide large improvements in speech intelligibility for CI users. These algorithms have been adopted from hearing aids and…

Cited by 0SourceScholar
2015

Interaural coherence preservation in MWF-based binaural noise reduction algorithms using partial noise estimation

ICASSP 2015accepted

Besides noise reduction an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of both desired and undesired sound sources. Recently an extension of the binaural Multi-channel Wiener filter (MWF), namely the MWF-IC, has been presented which aims to…

Cited by 0SourceScholar
2015

Joint acoustic and spectral modeling for speech dereverberation using non-negative representations

ICASSP 2015accepted

This paper proposes a single-channel speech dereverberation method enhancing the spectrum of the reverberant speech signal. The proposed method uses a non-negative approximation of the convolutive transfer function (N-CTF) to simultaneously estimate the magnitude spectrograms of the speech signal an…

Cited by 0SourceScholar
2015

Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparison

ICASSP 2015accepted

In this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities…

Cited by 0SourceScholar
2015

Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximation

ICASSP 2015accepted

In many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transf…

Cited by 14SourceScholar
2015

On application of non-negative matrix factorization for ad hoc microphone array calibration from incomplete noisy distances

ICASSP 2015accepted

We propose to use non-negative matrix factorization (NMF) to estimate the unknown pairwise distances and reconstruct a distance matrix for microphone array position calibration. We develop new multiplicative update rules for NMF with incomplete input matrix that take into account the symmetry of the…

Cited by 0SourceScholar
2015

Speaker change detection and speaker diarization using spatial information

ICASSP 2015accepted

In this paper, we present a novel speaker change detection and speaker diarization algorithm using spatial information in the form of features derived from estimated Room Impulse Response (RIR)s. A blind system identification approach is used to obtain an estimate of the RIRs, from which the C5 feat…

Cited by 13SourceScholar