← Search

Sharon Gannot

37 accepted papers

2026

DIFFUSION-BASED UNSUPERVISED AUDIO-VISUAL SPEECH SEPARATION IN NOISY ENVIRONMENTS WITH NOISE PRIOR

ICASSP 2026poster

In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise components, training exclusively on these individual signals rather than noisy mix…

Cited by 0SourcePDFScholar
2026

SPECTRAL OR SPATIAL? LEVERAGING BOTH FOR SPEAKER EXTRACTION IN CHALLENGING DATA CONDITIONS

ICASSP 2026poster

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker, our method integrates both sources of information to enhanc…

Cited by 0SourcePDFScholar
2025

Multi-Microphone Speech Emotion Recognition Using the Hierarchical Token-Semantic Audio Transformer Architecture

ICASSP 2025accepted

The performance of most emotion recognition systems degrades in real-life situations ("in the wild" scenarios) where the audio is contaminated by reverberation. Our study explores new methods to alleviate the performance degradation of Speech Emotion Recognition (SER) algorithms and develop a more r…

Cited by 0SourceScholar
2024

Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple Speakers

ICASSP 2024accepted

To estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different frequencies can be combined. According to how the SPS are combi…

Cited by 0SourceScholar
2024

LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading

ICLR 2024poster

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for challenging and realistic datasets such as LRS3. In this work…

2024

Unsupervised Acoustic Scene Mapping Based on Acoustic Features and Dimensionality Reduction

ICASSP 2024accepted

Classical methods for acoustic scene mapping require the estimation of the time difference of arrival (TDOA) between microphones. Unfortunately, TDOA estimation is very sensitive to reverberation and additive noise. We introduce an unsupervised data-driven approach that exploits the natural structur…

Cited by 0SourceScholar
2023

Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural Network

ICASSP 2023accepted

The interpretation and explanation of decision-making processes of neural networks are becoming a key factor in the deep learning field. Although several approaches have been presented for classification problems, the application to regression models needs to be further investigated. In this manuscr…

Cited by 0SourceScholar
2022

Closed-Form Single Source Direction-of-Arrival Estimator Using First-Order Relative Harmonic Coefficients

ICASSP 2022accepted

The relative harmonic coefficients (RHC), recently introduced as a multi-microphone spatial feature, demonstrates promising performance when applied to direction-of-arrival (DOA) estimation. All existing RHC-based DOA estimators suffer from a resolution limitation due to the inherent grid-based sear…

Cited by 0SourceScholar
2022

Low Resources Online Single-Microphone Speech Enhancement with Harmonic Emphasis

ICASSP 2022accepted

In this paper, we propose a deep neural network (DNN)-based single-microphone speech enhancement algorithm characterized by a short latency and low computational resources. Many speech enhancement algorithms suffer from low noise reduction capabilities between pitch harmonics, and in severe cases, t…

Cited by 0SourceScholar
2021

Evaluation and Comparison of Three Source Direction-of-Arrival Estimators Using Relative Harmonic Coefficients

ICASSP 2021accepted

A spherical harmonics domain source feature called relative harmonic coefficients (RHC) has recently been applied to address the source direction-of-arrival (DOA) estimation problem. This paper presents a compact evaluation and comparison between two existing RHC based DOA estimators: (i) a method u…

Cited by 0SourceScholar
2021

Misalignment Recognition in Acoustic Sensor Networks Using a Semi-Supervised Source Estimation Method and Markov Random Fields

ICASSP 2021accepted

In this paper, we consider the problem of acoustic source localization by acoustic sensor networks (ASNs) using a promising, learning-based technique that adapts to the acoustic environment. In particular, we look at the scenario when a node in the ASN is displaced from its position during training.…

Cited by 0SourceScholar
2021

Speech Enhancement with Mixture of Deep Experts with Clean Clustering Pre-Training

ICASSP 2021accepted

In this study we present a mixture of deep experts (MoDE) neural-network architecture for single microphone speech enhancement. Our architecture comprises a set of deep neural networks (DNNs), each of which is an ‘expert’ in a different speech spectral pattern such as phoneme. A gating DNN is respon…

Cited by 0SourceScholar
2020

A Composite DNN Architecture for Speech Enhancement

ICASSP 2020accepted

In speech enhancement, the use of supervised algorithms in the form of deep neural networks (DNNs) has become tremendously popular in recent years. The target function of the DNN (and the associated estimators) is often either a masking function applied to the noisy spectrum, or the clean log-spectr…

Cited by 0SourceScholar
2020

Low Complexity NLMS for Multiple Loudspeaker Acoustic ECHO Canceller Using Relative Loudspeaker Transfer Functions

ICASSP 2020accepted

Speech signals captured by a microphone mounted to a smart soundbar or speaker are inherently contaminated by echos. Modern smart devices are usually characterized by low computational capabilities and low memory resources; in these cases, a low-complexity acoustic echo canceller (AEC) may be prefer…

Cited by 0SourceScholar
2020

Maximum Likelihood Multi-Speaker Direction of Arrival Estimation Utilizing a Weighted Histogram

ICASSP 2020accepted

In this contribution, a novel maximum likelihood (ML) based direction of arrival (DOA) estimator for concurrent speakers in a noisy reverberant environment is presented. The DOA estimation task is formulated in the short-time Fourier transform (STFT) in two stages. In the first stage, a single local…

Cited by 0SourceScholar
2020

Unsupervised Multiple Source Localization Using Relative Harmonic Coefficients

ICASSP 2020accepted

This paper presents an unsupervised multi-source localization algorithm using a recently introduced feature called the relative harmonic coefficients. We derive a closed-form expression of the feature and briefly summarize its unique properties. We then exploit this feature to develop a single-sourc…

Cited by 0SourceScholar
2019

An Online Multiple-speaker DOA Tracking Using the CappÉ-Moulines Recursive Expectation-maximization Algorithm

ICASSP 2019accepted

In this paper, we present a multiple-speaker direction of arrival (DOA) tracking algorithm with a microphone array that utilizes the recursive EM (REM) algorithm proposed by Cappé and Moulines. In our model, all sources can be located in one of a predefined set of candidate DOAs. Accordingly, the re…

Cited by 0SourceScholar
2019

Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and Diarization

ICASSP 2019accepted

This paper investigates localization of an arbitrary number of simultaneously active speakers in an acoustic enclosure. We propose an algorithm capable of estimating the number of speakers, using reliability information to obtain robust estimation results in adverse acoustic scenarios and estimating…

Cited by 8SourceScholar
2018

DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming

ICASSP 2018accepted

In this paper, we present a new control mechanism for LCMV beamforming. Application of the LCMV beamformer to speaker separation tasks requires accurate estimates of its building blocks, e.g. the noise spatial cross-power spectral density (cPSD) matrix and the relative transfer function (RTF) of all…

Cited by 0SourceScholar
2018

Multi-View Source Localization Based on Power Ratios

ICASSP 2018accepted

Despite attracting significant research efforts, the problem of source localization in noisy and reverberant environments remains challenging. Novel learning-based methods attempt to solve the problem by modelling the acoustic environment from the observed data. Typically, appropriate feature vector…

Cited by 0SourceScholar
2017

An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixtures

ICASSP 2017accepted

We present a probabilistic model for joint source separation and diarisation of multichannel convolutive speech mixtures. We build upon the framework of local Gaussian model (LGM) with non-negative matrix factorization (NMF). The diarisation is introduced as a temporal labeling of each source in the…

Cited by 0SourceScholar
2017

Comparison of two binaural beamforming approaches for hearing aids

ICASSP 2017accepted

Beamforming algorithms in binaural hearing aids are crucial to improve speech understanding in background noise for hearing impaired persons. In this study, we compare and evaluate the performance of two recently proposed minimum variance (MV) beamforming approaches for binaural hearing aids. The bi…

Cited by 0SourceScholar
2017

Source tracking using moving microphone arrays for robot audition

ICASSP 2017accepted

Intuitive spoken dialogues are a prerequisite for human-robot interaction. In many practical situations, robots must be able to identify and focus on sources of interest in the presence of interfering speakers. Techniques such as spatial filtering and blind source separation are therefore often used…

Cited by 0SourceScholar
2016

An inverse-gamma source variance prior with factorized parameterization for audio source separation

ICASSP 2016accepted

In this paper we present a new statistical model for the power spectral density (PSD) of an audio signal and its application to multichannel audio source separation (MASS). The source signal is modeled with the local Gaussian model (LGM) and we propose to model its variance with an inverse-Gamma dis…

Cited by 0SourceScholar
2016

Extensions of the binaural MWF with interference reduction preserving the binaural cues of the interfering source

ICASSP 2016accepted

Recently, an extension of the binaural multichannel Wiener filter (BMWF), referred to as BMWF-IRo, was presented in which an interference rejection constraint was added to the BMWF cost function. Although the BMWF-IRo aims to entirely suppress the interfering source, residual interfering sources (as…

Cited by 0SourceScholar
2016

Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aids

ICASSP 2016accepted

Besides noise reduction, an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. For the desired speech source and an interfering source, e.g., competing speaker, this can be achieved by preserving their relative transfer functi…

Cited by 0SourceScholar
2016

Joint maximum likelihood estimation of late reverberant and speech power spectral density in noisy environments

ICASSP 2016accepted

An estimate of the power spectral density (PSD) of the late reverberation is often required by dereverberation algorithms. In this work, we derive a novel multichannel maximum likelihood (ML) estimator for the PSD of the reverberation that can be applied in noisy environments. Since the anechoic spe…

Cited by 0SourceScholar
2016

Manifold-based Bayesian inference for semi-supervised source localization

ICASSP 2016accepted

Sound source localization is addressed by a novel Bayesian approach using a data-driven geometric model. The goal is to recover the target function that attaches each acoustic sample, formed by the measured signals, with its corresponding position. The estimation is derived by maximizing the posteri…

Cited by 0SourceScholar
2016

Non-stationary noise power spectral density estimation based on regional statistics

ICASSP 2016accepted

Estimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past a…

Cited by 0SourceScholar
2015

Binaural multichannel Wiener filter with directional interference rejection

ICASSP 2015accepted

In this paper we consider an acoustic scenario with a desired source and a directional interference picked up by hearing devices in a noisy and reverberant environment. We present an extension of the binaural multichannel Wiener filter (BMWF), by adding an interference rejection constraint to its co…

Cited by 0SourceScholar
2015

Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtraction

ICASSP 2015accepted

This paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresp…

Cited by 0SourceScholar
2015

Nested generalized sidelobe canceller for joint dereverberation and noise reduction

ICASSP 2015accepted

Speech signal is often contaminated by both room reverberation and ambient noise. In this contribution, we propose a nested generalized sidelobe canceller (GSC) beamforming structure, comprising an inner and an outer GSC beamformers (BFs), that decouple the speech dereverberation and the noise reduc…

Cited by 0SourceScholar
2015

Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whitening method

ICASSP 2015accepted

Microphone array processing utilize spatial separation between the desired speaker and interference signal for speech enhancement. The transfer functions (TFs) relating the speaker component at a reference microphone with all other microphones, denoted as the relative TFs (RTFs), play an important r…

Cited by 0SourceScholar