← Search

Zhongweiyang Xu

6 accepted papers

2026

ARRAYDPS-REFINE: GENERATIVE REFINEMENT OF DISCRIMINATIVE MULTI-CHANNEL SPEECH ENHANCEMENT

ICASSP 2026poster

Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives, especially under challenging environmental noise conditions. Inspir…

Cited by 0SourcePDFScholar
2026

Contrastive Diffusion Guidance for Spatial Inverse Problems

ICLR 2026poster

We consider the inverse problem of reconstructing the spatial layout of a place, a home floorplan for example, from a user’s movements inside that layout. Direct inversion is ill-posed since many floorplans can explain the same movement trajectories. We adopt a diffusion-based posterior sampler to…

Cited by 0SourceScholar
2025

ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior

ICML 2025poster

Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are all unknown.…

2024

SPATIALCODEC: Neural Spatial Speech Coding

ICASSP 2024accepted

In this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We propose a neural spatial audio coding framework that achieves…

Cited by 0SourceScholar
2024

uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models

ICASSP 2024accepted

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional…

Cited by 15SourceScholar
2023

Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction

ICASSP 2023accepted

Audiovisual target speaker extraction is the task of separating, from an audio mixture, the speaker whose face is visible in an accompanying video. Published approaches typically upsample the video or downsample the audio, then fuse the two streams using concatenation, multiplication, or cross-modal…

Cited by 0SourceScholar