← Search

Xiguang Zheng

7 accepted papers

2025

A Metric for Predicting the Quality of Ambisonic Spatial Audio Reproduced Using Spatially Interpolated or Extrapolated Room Impulse Responses

ICASSP 2025accepted

In virtual reality (VR), sound sources are convolved with room impulse responses (RIRs) to create immersive and dynamic audio experiences. Assessing the quality of spatial audio synthesis in VR is challenging. Subjective listening tests are accurate, but they are time-consuming and costly. This pape…

Cited by 0SourceScholar
2024

BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution

ICASSP 2024accepted

Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuat…

Cited by 0SourceScholar
2023

A Low-Latency Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation

ICASSP 2023accepted

This paper describes our submission to the fourth Acoustic Echo Cancellation (AEC) Challenge, which is part of ICASSP 2023 Signal Processing Grand Challenge. The proposed system is developed based on our earlier system submitted to the ICASSP 2022 AEC challenge with significant latency and network s…

Cited by 0SourceScholar
2022

A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation

ICASSP 2022accepted

Deep learning based wideband (16kHz) acoustic echo cancellation (AEC) approaches have surpassed traditional methods. This work proposes a deep hierarchical fusion (DHF) network with intra-network and inter-network fusion to further improve the wideband AEC performance. Meanwhile, this work extends t…

Cited by 0SourceScholar
2022

A Two-Step Backward Compatible Fullband Speech Enhancement System

ICASSP 2022accepted

Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system is proposed in this paper. Compared to the existing full-band systems that util…

Cited by 0SourceScholar
2022

L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment

ICASSP 2022accepted

The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition <sup xmlns:mml="http://www.w3.org/1998/Math…

Cited by 0SourceScholar
2022

Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement

ICASSP 2022accepted

Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, perso…

Cited by 0SourceScholar