← Search

Emanuël A. P. Habets

31 accepted papers

2026

MATCHING REVERBERANT SPEECH THROUGH LEARNED ACOUSTIC EMBEDDINGS AND FEEDBACK DELAY NETWORKS

ICASSP 2026oral

Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key challenge, especially when explicit acoustic measurements are unava…

Cited by 0SourcePDFScholar
2025

GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning

ICASSP 2025accepted

Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discri…

Cited by 0SourceScholar
2025

Low-Complexity Neural Speech Dereverberation With Adaptive Target Control

ICASSP 2025accepted

Existing neural network-based speech dereverberation approaches use a fixed-length early reflection part of the reverberant signal as the target for estimation, irrespective of the severity of reverberation. Such an approach often leads to distortions in the enhanced signals in highly reverberant sc…

Cited by 0SourceScholar
2025

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

ICASSP 2025accepted

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers. The speech quality across the speaker set typically is diverse…

Cited by 0SourceScholar
2025

On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs

ICASSP 2025accepted

Neural audio signal codecs have attracted significant attention in recent years. In essence, the impressive low bitrate achieved by such encoders is enabled by learning an abstract representation that captures the properties of encoded signals, e.g., speech. In this work, we investigate the relation…

Cited by 0SourceScholar
2024

Binaural Rendering of Heterogeneous Sound Sources with Extent

ICASSP 2024accepted

In spatial audio rendering applications, it is often desired to render sound sources with a certain spatial extent in a realistic way. While existing methods mainly consider rendering of homogeneously extended sound sources (i.e., with constant radiation characteristics over the extent), rendering o…

Cited by 0SourceScholar
2024

Odaq: Open Dataset of Audio Quality

ICASSP 2024accepted

Research into the prediction and analysis of perceived audio quality is hampered by the scarcity of openly available datasets of audio signals accompanied by corresponding subjective quality scores. To address this problem, we present the Open Dataset of Audio Quality (ODAQ), a new dataset containin…

Cited by 0SourceScholar
2024

Sector-Based Interference Cancellation for Robust Keyword Spotting Applications Using an Informed MPDR Beamformer

ICASSP 2024accepted

A low-complexity, sector-based interference cancellation approach is proposed for voice-controlled devices, e.g., smart speakers. We propose an informed minimum power distortionless response beamformer that provides an optimal trade-off between noise reduction, dereverberation, and interference canc…

Cited by 0SourceScholar
2023

Better Together: Dialogue Separation and Voice Activity Detection for Audio Personalization in TV

ICASSP 2023accepted

In TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estimate them to enable personalization. DS was shown to provide clear benefits for th…

Cited by 0SourceScholar
2023

Contrastive Representation Learning for Acoustic Parameter Estimation

ICASSP 2023accepted

A study is presented in which a contrastive learning approach is used to extract low-dimensional representations of the acoustic environment from single-channel, reverberant speech signals. Convolution of room impulse responses (RIRs) with anechoic source signals is leveraged as a data augmentation…

Cited by 0SourceScholar
2023

Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech Synthesis

ICASSP 2023accepted

In recent years, the quality of text-to-speech (TTS) synthesis vastly improved due to deep-learning techniques, with parallel architectures, in particular, providing excellent synthesis quality at fast inference. Training these models usually requires speech recordings, corresponding phoneme-level t…

Cited by 0SourceScholar
2023

Multi-Microphone Speaker Separation by Spatial Regions

ICASSP 2023accepted

We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual spatial regions as captured by a reference microphone while re…

Cited by 0SourceScholar
2022

Blind Reverberation Time Estimation in Dynamic Acoustic Conditions

ICASSP 2022accepted

The estimation of reverberation time from real-world signals plays a central role in a wide range of applications. In many scenarios, acoustic conditions change over time which in turn requires the estimate to be updated continuously. Previously proposed methods involving deep neural networks were m…

Cited by 0SourceScholar
2021

Direction Preserving Wind Noise Reduction Of B-Format Signals

ICASSP 2021accepted

Noise reduction in B-format recordings is particularly challenging as it concurrently requires to suppress undesired signals and preserve the spatial properties of the acoustic environment. In particular, wind noise poses an undesirable acoustic condition outdoors. In this work, methods to reduce wi…

Cited by 0SourceScholar
2021

Efficient Training Data Generation for Phase-Based DOA Estimation

ICASSP 2021accepted

Deep learning (DL) based direction of arrival (DOA) estimation is an active research topic and currently represents the state-of-the-art. Usually, DL-based DOA estimators are trained with recorded data or computationally expensive generated data. Both data types require significant storage and exces…

Cited by 0SourceScholar
2020

Data-Driven Wind Speed Estimation Using Multiple Microphones

ICASSP 2020accepted

A deep neural network (DNN) based approach for estimating the speed of airflows using closely-spaced microphones is proposed. The spatial characteristics of wind noise measured with a smallaperture array are exploited, i.e., the low-frequency spatial coherence of wind noise signals is used as an inp…

Cited by 0SourceScholar
2020

Low Complexity NLMS for Multiple Loudspeaker Acoustic ECHO Canceller Using Relative Loudspeaker Transfer Functions

ICASSP 2020accepted

Speech signals captured by a microphone mounted to a smart soundbar or speaker are inherently contaminated by echos. Modern smart devices are usually characterized by low computational capabilities and low memory resources; in these cases, a low-complexity acoustic echo canceller (AEC) may be prefer…

Cited by 0SourceScholar
2018

Classification vs. Regression in Supervised Learning for Single Channel Speaker Count Estimation

ICASSP 2018accepted

The task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene classification. Building upon powerful machine learning methodol…

Cited by 0SourceScholar
2018

Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio Estimation

ICASSP 2018accepted

Non-intrusive estimators for acoustic parameters like the direct-to-reverberation ratio (DRR) are useful tools but still perform weakly as shown in the acoustic characterization of environments (ACE) challenge. In this paper, we develop a novel dual-channel metric based on the modulation energy doma…

Cited by 0SourceScholar
2017

Time of arrival disambiguation using the linear Radon transform

ICASSP 2017accepted

Echo labeling, the challenging task of assigning acoustic reflections to image sources, is equivalent to the highly-important disambiguation task in room geometry inference. A method using the Radon transform, an image processing tool, is proposed to address this challenge. The method relies on acou…

Cited by 0SourceScholar
2016

A low complexity weighted least squares narrowband DOA estimator for arbitrary array geometries

ICASSP 2016accepted

An increasing number of spatial filtering approaches requires narrowband direction-of-arrival (DOA) estimates. State-of-the-art (SOA) estimators such as root-MUSIC and ESPRIT are computationally complex and can be used only with specific array geometries. In this work, a low complexity DOA estimator…

Cited by 0SourceScholar
2016

Conditional MMSE-based single-channel speech enhancement using inter-frame and inter-band correlations

ICASSP 2016accepted

Obtaining an estimate of clean speech for each time-frequency (TF) unit continues to be of importance in single-channel speech enhancement. Recently, it has been proposed to exploit inter-frame and interband correlations in a variety of speech processing applications. To estimate the clean speech, w…

Cited by 0SourceScholar
2016

Insight into a phase modulation technique for signal decorrelation in multi-channel acoustic echo cancellation

ICASSP 2016accepted

High coherence between the loudspeaker signals in a multichannel communication set-up is known to be detrimental to the performance of a multi-channel acoustic echo cancellation (MC-AEC) system. The MC-AEC performance can be improved by decorrelating the loudspeaker signals prior to their reproducti…

Cited by 0SourceScholar
2016

Joint maximum likelihood estimation of late reverberant and speech power spectral density in noisy environments

ICASSP 2016accepted

An estimate of the power spectral density (PSD) of the late reverberation is often required by dereverberation algorithms. In this work, we derive a novel multichannel maximum likelihood (ML) estimator for the PSD of the reverberation that can be applied in noisy environments. Since the anechoic spe…

Cited by 0SourceScholar
2015

A Bayesian approach to spatial filtering and diffuse power estimation for joint dereverberation and noise reduction

ICASSP 2015accepted

A spatial filter, with L linear constraints that are based on instantaneous narrowband direction-of-arrival (DOA) estimates, was recently proposed to obtain a desired spatial response for at most L sound sources. In noisy and reverberant environments, it becomes difficult to get reliable instantaneo…

Cited by 0SourceScholar
2015

A state-space partitioned-block adaptive filter for echo cancellation using inter-band correlations in the Kalman gain computation

ICASSP 2015accepted

A partitioned-block-based architecture for a model-based acoustic echo canceller in the frequency domain was recently presented. Partitioned-block-based frequency domain adaptive filters provide a lower algorithmic delay compared to the non-partitioned formulations, which is achieved by partitioning…

Cited by 0SourceScholar
2015

Direct-ambient decomposition using parametric wiener filtering with spatial cue control

ICASSP 2015accepted

A method for decomposing audio signals into direct signals and ambient signals is described that can be applied to sound post-production and reproduction. The proposed method is based on a parametric multichannel Wiener filter (MWF) that enables a trade-off between the attenuation of the interfering…

Cited by 0SourceScholar
2015

Minimum Bayes risk signal detection for speech enhancement based on a narrowband DOA model

ICASSP 2015accepted

A desired speech signal in hands-free communication systems is often degraded by background noise and interferers. Data-dependent spatial filters for desired speech extraction depend on the power spectral density (PSD) matrices of the desired and the undesired signals, which are commonly estimated r…

Cited by 0SourceScholar
2015

Nested generalized sidelobe canceller for joint dereverberation and noise reduction

ICASSP 2015accepted

Speech signal is often contaminated by both room reverberation and ambient noise. In this contribution, we propose a nested generalized sidelobe canceller (GSC) beamforming structure, comprising an inner and an outer GSC beamformers (BFs), that decouple the speech dereverberation and the noise reduc…

Cited by 0SourceScholar
2015

Residual noise control using a parametric multichannel Wiener filter

ICASSP 2015accepted

Multichannel noise reduction techniques are commonly used in speech communication applications. In these applications, it is often desired to maintain a residual amount of background noise to avoid perceptually unpleasant artifacts, such as musical tones or time periods of complete silence. Noise re…

Cited by 0SourceScholar