← Search

Zhong-Qiu Wang

30 accepted papers

2026

FORWARD CONVOLUTIVE PREDICTION FOR FRAME ONLINE MONAURAL SPEECH DEREVERBERATION BASED ON KRONECKER PRODUCT DECOMPOSITION

ICASSP 2026poster

Dereverberation has long been a crucial research topic in speech processing, aiming to alleviate the adverse effects of reverberation in voice communication and speech interaction systems. Among existing approaches, forward convolutional prediction (FCP) has recently attracted attention. It typicall…

Cited by 0SourcePDFScholar
2026

MC-LExt: Multi-Channel Target Speaker Extraction with Onset-Prompted Speaker Conditioning Mechanism

ICASSP 2026poster

Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or speaker embeddings. However, DOA-based approaches depend on expl…

Cited by 0SourcePDFScholar
2026

Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR

ICASSP 2026oral

In multi-channel speech enhancement and robust automatic speech recognition (ASR), beamforming can typically improve the signal-to-noise ratio (SNR) of the target speaker and produce reliable enhancement with little distortion to target speech. With this observation, we propose to leverage beamforme…

Cited by 0SourcePDFScholar
2025

30+ Years of Source Separation Research: Achievements and Future Challenges

ICASSP 2025accepted

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> anniversary, we review the major contribut…

Cited by 0SourceScholar
2025

ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior

ICML 2025poster

Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are all unknown.…

2024

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

ICASSP 2024accepted

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompte…

Cited by 0SourceScholar
2023

FNeural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated full- and sub-band Modeling

ICASSP 2023accepted

We propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in the short-time Fourier transform (STFT) domain. The model maintains an information highway to flow an over-complete inp…

Cited by 0SourceScholar
2023

Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand Challenge

ICASSP 2023accepted

In this work we detail our submission to the Clarity ICASSP 2023 grand challenge, in which participants have to develop a strong target speech enhancement system for hearing-aid (HA) devices in noisy-reverberant environments. Our system builds on our previous submission at the Second Clarity Enhance…

Cited by 0SourceScholar
2023

TF-GRIDNET: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation

ICASSP 2023accepted

We propose TF-GridNet, a novel multi-path deep neural network (DNN) operating in the time-frequency (T-F) domain, for monaural talker-independent speaker separation in anechoic conditions. The model stacks several multi-path blocks, each consisting of an intra-frame spectral module, a sub-band tempo…

Cited by 0SourceScholar
2023

UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training Mixtures

NeurIPS 2023poster

In reverberant conditions with multiple concurrent speakers, each microphone acquires a mixture signal of multiple speakers at a different location. In over-determined conditions where the microphones out-number speakers, we can narrow down the solutions to speaker images and realize unsupervised sp…

Cited by 13SourcePDFScholar
2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

ICASSP 2022accepted

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech synthesis, they are still lagging behind in speech enhancement. This work leverage…

Cited by 0SourceScholar
2022

Locate This, Not that: Class-Conditioned Sound Event DOA Estimation

ICASSP 2022accepted

Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propose an alternative class-conditioned SELD model for situations where we may not be interested in localizing all classes al…

Cited by 0SourceScholar
2022

The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks

ICASSP 2022accepted

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from noise, speech from speech, musical instruments from each other, or sound events fr…

Cited by 0SourceScholar
2022

Towards Low-Distortion Multi-Channel Speech Enhancement: The ESPNET-Se Submission to the L3DAS22 Challenge

ICASSP 2022accepted

This paper describes our submission to the L3DAS22 Challenge Task 1, which consists of speech enhancement with 3D Ambisonic microphones. The core of our approach combines Deep Neural Network (DNN) driven complex spectral mapping with linear beamformers such as the multi-frame multi-channel Wiener fi…

Cited by 0SourceScholar
2019

Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective

ICASSP 2019accepted

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture oftwo sources, with their magnitudes accurately estimated and under a geometric constraint…

Cited by 0SourceScholar
2018

Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASR

ICASSP 2018accepted

Deep learning based single-channel time-frequency (T-F) masking has shown considerable potential for beamforming and robust ASR. This paper proposes a simple but novel relative transfer function (RTF) estimation algorithm for microphone arrays, where the RTF between a reference signal and a non-refe…

Cited by 0SourceScholar
2018

Multi-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation

ICASSP 2018accepted

The recently-proposed deep clustering algorithm represents a fundamental advance towards solving the cocktail party problem in the single-channel case. When multiple microphones are available, spatial information can be leveraged to differentiate signals from different directions. This study combine…

Cited by 0SourceScholar
2018

On Spatial Features for Supervised Speech Separation and its Application to Beamforming and Robust ASR

ICASSP 2018accepted

This study integrates complementary spectral and spatial information to elevate deep learning based time-frequency masking and acoustic beamforming. Coherence and directional features are designed as additional input features for deep neural network training to remove diffuse noise and other directi…

Cited by 0SourceScholar
2017

A speech enhancement algorithm by iterating single- and multi-microphone processing and its application to robust ASR

ICASSP 2017accepted

We propose a speech enhancement algorithm based on single- and multi-microphone processing techniques. The core of the algorithm estimates a time-frequency mask which represents the target speech and use masking-based beamforming to enhance corrupted speech. Specifically, in single-microphone proces…

Cited by 0SourceScholar
2017

Learning utterance-level representations for speech emotion and age/gender recognition using deep neural networks

ICASSP 2017accepted

Accurately recognizing speaker emotion and age/gender from speech can provide better user experience for many spoken dialogue systems. In this study, we propose to use deep neural networks (DNNs) to encode each utterance into a fixed-length vector by pooling the activations of the last hidden layer…

Cited by 0SourceScholar