← Search

Buye Xu

18 accepted papers

2025

Advancing Active Speaker Detection for Egocentric Videos

ICASSP 2025accepted

This paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing th…

Cited by 0SourceScholar
2025

Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement

ICASSP 2025accepted

Deep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specificall…

Cited by 0SourceScholar
2025

Reexamining the Efficacy of MetricGAN for Speech Enhancement

ICASSP 2025accepted

MetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluat…

Cited by 0SourceScholar
2025

Robust Frame-level Speaker Localization in Reverberant and Noisy Environments by Exploiting Phase Difference Losses

ICASSP 2025accepted

This paper investigates robust speaker localization at the frame level on the basis of complex spectral mapping, which is capable of learning both the magnitude and phase of the target signal. Unlike prevailing deep learning methods for speaker localization, we perform MIMO (multi-input multi-output…

Cited by 0SourceScholar
2024

A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech Enhancement

ICASSP 2024accepted

Self-supervised learned models have been found to be very effective for tasks such as automatic speech recognition, speaker identification, and others. However, their utility in speech enhancement systems is yet to be firmly established, and perhaps slightly misunderstood. In this paper, we investig…

Cited by 0SourceScholar
2024

Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain

ICASSP 2024accepted

We introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual…

Cited by 0SourceScholar
2024

Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement

ICASSP 2024accepted

We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and temporal processing within deep neural network (DNN) layers.…

Cited by 0SourceScholar
2024

Leveraging Sound Localization to Improve Continuous Speaker Separation

ICASSP 2024accepted

Continuous speaker separation aims to separate overlapping speakers in real-world environments like meetings, but it often falls short in isolating speech segments of a single speaker. This leads to split signals that adversely affect downstream applications such as automatic speech recognition and…

Cited by 9SourceScholar
2024

On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech Enhancement

ICASSP 2024accepted

We introduce a time-domain framework for efficient multichannel speech enhancement, emphasizing low latency and computational efficiency. This framework incorporates two compact deep neural networks (DNNs) surrounding a multichannel neural Wiener filter (NWF). The first DNN enhances the speech signa…

Cited by 0SourceScholar
2023

LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders

ICASSP 2023accepted

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker’s lip movements. This approach has been shown to yield improvements over audio-only speech enhancement, particularly for the removal of interferin…

Cited by 0SourceScholar
2023

Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech Enhancement

ICASSP 2023accepted

Most speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivariate Gaussian negative log-likelihood (NLL) improves SE performance at no extra c…

Cited by 2SourceScholar
2023

Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio

ICASSP 2023accepted

Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a s…

Cited by 124SourceScholar
2022

Continual Self-Training With Bootstrapped Remixing For Speech Enhancement

ICASSP 2022accepted

We propose RemixIT, a simple and novel self-supervised training method for speech enhancement. The proposed method is based on a continuously self-training scheme that overcomes limitations from previous studies including assumptions for the in-domain noise distribution and having access to clean ta…

Cited by 0SourceScholar
2022

Multichannel Speech Enhancement Without Beamforming

ICASSP 2022accepted

Deep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post…

Cited by 24SourceScholar
2022

TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement

ICASSP 2022accepted

In this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes sp…

Cited by 0SourceScholar
2021

NORESQA: A Framework for Speech Quality Assessment using Non-Matching References

NeurIPS 2021poster

The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA. Clearly, these methods fail in real-world scenarios where the grou…

2018

Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term Memory

ICASSP 2018accepted

Human speech is usually distorted by room reverberation. These corruptions degrade speech quality and intelligibility, especially under a long reverberation time, and they also pose a serious problem for many speech-related applications such as automatic speech recognition. In this paper, we propose…

Cited by 0SourceScholar