← Search

Shlomo E. Chazan

9 accepted papers

2026

SPECTRAL OR SPATIAL? LEVERAGING BOTH FOR SPEAKER EXTRACTION IN CHALLENGING DATA CONDITIONS

ICASSP 2026poster

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker, our method integrates both sources of information to enhanc…

Cited by 0SourcePDFScholar
2025

Automatic Detection of Domain Shifts in Speech Enhancement Systems Using Confidence-Based Metrics

ICASSP 2025accepted

Introducing a domain shift, such as a change in language or environment, to a well-trained speech enhancement system can cause severe performance degradation. Most current research assumes that a domain shift has already been detected and focuses on either supervised or unsupervised domain adaptatio…

Cited by 0SourceScholar
2024

C-CLAPA: Improving Text-Audio Cross Domain Retrieval with Captioning and Augmentations

ICASSP 2024accepted

In this paper, we introduce Captioning decoder Contrastive Language-Audio Pretraining with data Augmantation (C-CLAPA), a new Audio-Text model for the Cross Domain Retrieval (CDR) task. The model’s backbone is comprised of two encoders, one for the text and the other for the audio. The embedding vec…

Cited by 0SourceScholar
2021

Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy Settings

ICASSP 2021accepted

We present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the time domain, together with parameter sharing between all sep…

Cited by 0SourceScholar
2021

Speech Enhancement with Mixture of Deep Experts with Clean Clustering Pre-Training

ICASSP 2021accepted

In this study we present a mixture of deep experts (MoDE) neural-network architecture for single microphone speech enhancement. Our architecture comprises a set of deep neural networks (DNNs), each of which is an ‘expert’ in a different speech spectral pattern such as phoneme. A gating DNN is respon…

Cited by 0SourceScholar
2020

A Composite DNN Architecture for Speech Enhancement

ICASSP 2020accepted

In speech enhancement, the use of supervised algorithms in the form of deep neural networks (DNNs) has become tremendously popular in recent years. The target function of the DNN (and the associated estimators) is often either a masking function applied to the noisy spectrum, or the clean log-spectr…

Cited by 0SourceScholar
2018

DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming

ICASSP 2018accepted

In this paper, we present a new control mechanism for LCMV beamforming. Application of the LCMV beamformer to speaker separation tasks requires accurate estimates of its building blocks, e.g. the noise spatial cross-power spectral density (cPSD) matrix and the relative transfer function (RTF) of all…

Cited by 0SourceScholar