← Search

Changbao Zhu

6 accepted papers

2026

Rethinking Flow and Diffusion Bridge Models for Speech Enhancement

AAAI 2026technical

Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a fra

Cited by 0SourcePDFScholar
2024

A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity Detection

ICASSP 2024accepted

Although deep learning (DL) based end-to-end models have shown outstanding performance in multi-channel speech extraction, their practical applications on edge devices are restricted due to their high computational complexity. In this paper, we propose a hybrid system that can more effectively integ…

Cited by 0SourceScholar
2024

GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources

ICASSP 2024accepted

While modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this pape…

Cited by 0SourceScholar
2023

Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech Enhancement

ICASSP 2023accepted

MetricGAN and its variations have been proven to be an effective wide-band speech enhancement model. In this paper, we expand it to full-band enhancement by combining our recently proposed learnable spectral dimension compression mapping strategy. The encoder-decoder structure with a time-frequency…

Cited by 0SourceScholar
2023

TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty

ICASSP 2023accepted

In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which do…

Cited by 0SourceScholar
2015

Overview of the EVS codec architecture

ICASSP 2015accepted

The recently standardized 3GPP codec for Enhanced Voice Services (EVS) offers new features and improvements for low-delay real-time communication systems. Based on a novel, switched low-delay speech/audio codec, the EVS codec contains various tools for better compression efficiency and higher qualit…

Cited by 169SourceScholar