Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement
Zehua Zhang, Shiyun Xu, Xuyi Zhuang, Lianyu Zhou, Heng Li, Mingjiang Wang
Abstract
In denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a multi-axis gated multilayer perceptron to build global and local attention modules of linear complexity for extracting speech features. Furthermore, we use the residual channel attention block to further filter out important speech features. On the blind test dataset of the Deep Noise Suppression Challenge, our proposed TSUNet has a massive advantage over other state-of-the-art models. TSUNet performs significantly better than the most recent models at noisy-reverberant speech enhancement.
BibTeX
@inproceedings{icassp2023_twostageunetwith,
title = {Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement},
author = {Zehua Zhang and Shiyun Xu and Xuyi Zhuang and Lianyu Zhou and Heng Li and Mingjiang Wang},
booktitle = {ICASSP 2023},
year = {2023}
}