← Search

Chengshi Zheng

19 accepted papers

2026

GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks

AAAI 2026technical

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does

Cited by 0SourcePDFScholar
2026

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

AAAI 2026technical

Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in thi

Cited by 0SourcePDFScholar
2025

Audiogram-Informed End-to-End Noise Reduction and Wide Dynamic Range Compression for Hearing Aids

ICASSP 2025accepted

Wide dynamic range compression (WDRC) provides level-dependent amplification, intended to make the output of a hearing aid fall between the hearing threshold and the highest comfortable level of the listener. Hearing aids often combine noise reduction with WDRC, applied sequentially. Unfortunately,…

Cited by 0SourceScholar
2025

BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement

AAAI 2025technical

Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the…

Cited by 0SourcePDFScholar
2025

BridgeVoC: Neural Vocoder with Schrödinger Bridge

IJCAI 2025

While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restora

Cited by 0SourcePDFScholar
2025

DSINet: Towards Real-Time Target Speaker Extraction with Dynamic Speaker Information Fusion

ICASSP 2025accepted

Target speaker extraction (TSE) aims to directly extract the desired speech given enrollment utterances of the target speaker. Despite significant progress in recent years, most existing methods remain non-causal and computationally intensive. This paper introduces DSINet, a real-time time-frequency…

Cited by 0SourceScholar
2025

DeepPEM-AFC: An Improved Prediction-Error-Method-based Adaptive Feedback Cancellation with Deep Learning for Hearing Aids

ICASSP 2025accepted

Hearing assistive devices aim to compensate hearing loss for hearing-impaired listeners, and their maximum stable gain (MSG) is constrained because of the existence of the acoustic feedback between the receiver and microphone, resulting in their inefficiency for individuals with severe or profound h…

Cited by 0SourceScholar
2025

Learning Neural Vocoder from Range-Null Space Decomposition

IJCAI 2025

Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned chal

2024

All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays

ICASSP 2024accepted

Existing frame-wise neural beamformers for speech extraction can obtain promising performance in relatively high signal-to-noise ratio (SNR) scenarios using small microphone arrays, while they still suffer from performance degradation in relatively low SNR environments, e.g., SNR<-5 dB. As an attemp…

Cited by 0SourceScholar
2024

BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution

ICASSP 2024accepted

Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuat…

Cited by 0SourceScholar
2023

Gesper: A Unified Framework for General Speech Restoration

ICASSP 2023accepted

This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by spee…

Cited by 0SourceScholar
2022

Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement

ICASSP 2022accepted

Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance. Motivated by that, we propose a dual-branch attention-in-attention transformer dubbed DB-AIAT to handle both coarse…

Cited by 0SourceScholar
2022

Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement

ICASSP 2022accepted

Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and Beamforming, and two core modules are devised accordingly, namely EM and BM. For EM, instead of estimating spatial covariance matrix explicitly, the…

Cited by 0SourceScholar
2022

Joint Magnitude Estimation and Phase Recovery Using Cycle-In-Cycle GAN for Non-Parallel Speech Enhancement

ICASSP 2022accepted

For the lack of adequate paired noisy-clean speech corpus in many real scenarios, non-parallel training is a promising task for DNN-based speech enhancement methods. However, because of the severe mismatch between input and target speeches, many previous studies only focus on the magnitude spectrum…

Cited by 0SourceScholar
2022

Taylor, Can You Hear Me Now? A Taylor-Unfolding Framework for Monaural Speech Enhancement

IJCAI 2022poster

While the deep learning techniques promote the rapid development of the speech enhancement (SE) community, most schemes only pursue the performance in a black-box manner and lack adequate model interpretability. Inspired by Taylor's approximation theory, we propose an interpretable decoupling-style…

2021

ICASSP 2021 Acoustic Echo Cancellation Challenge: Integrated Adaptive Echo Cancellation with Time Alignment and Deep Learning-Based Residual Echo Plus Noise Suppression

ICASSP 2021accepted

This paper describes a three-stage acoustic echo cancellation (AEC) and suppression framework for the ICASSP 2021 AEC Challenge. In the first stage, a partitioned block frequency domain adaptive filtering is implemented to cancel the linear echo components without introducing the near-end speech dis…

Cited by 0SourceScholar
2021

ICASSP 2021 Deep Noise Suppression Challenge: Decoupling Magnitude and Phase Optimization with a Two-Stage Deep Network

ICASSP 2021accepted

It remains a tough challenge to recover the speech signals contaminated by various noises under real acoustic environments. To this end, we propose a novel system for denoising in the complicated applications, which is mainly comprised of two pipelines, namely a two-stage network and a post-processi…

Cited by 0SourceScholar