Hybrid Feature Global Attention Network for Noisy-reverberant Speech Enhancement
Zehua Zhang, Shiyun Xu, Yinghan Cao, Changjun He
Abstract
Deep neural network-based speech enhancement methods have become widespread, with one of its fundamental aspects being the effective extraction and application of features in the time-frequency domain. This paper proposes a hybrid feature global attention network (HFGANet) designed to efficiently extract time-frequency domain features. HFGANet incorporates a hybrid gated multilayer perceptron (HgMLP) that effectively captures local, global, and inter-window features in the time-frequency domain to create hybrid representations. In contrast to traditional convolutional recurrent neural network architectures, this paper innovatively proposes a global attention structure to leverage these hybrid features. The proposed global attention block enhances the integration of local and global features. Additionally, we introduce Temporal Mamba and Frequency Mamba to further improve the model's ability to capture contextual information in both time and frequency dimensions. On the 1st Deep Noise Suppression Challenge blind test set with reverberation, HFGANet achieves 3.51 WB-PESQ, 95.03% STOI, and 17.72 SI-SDR, while maintaining a lower parameter count compared to state-of-the-art models. In the task of noisy-reverberant speech enhancement, our model achieved an improvement of 1.22 in PESQ, 17.4% in STOI, and 1.63 in DNSMOS.
BibTeX
@inproceedings{icassp2025_hybridfeatureglo,
title = {Hybrid Feature Global Attention Network for Noisy-reverberant Speech Enhancement},
author = {Zehua Zhang and Shiyun Xu and Yinghan Cao and Changjun He},
booktitle = {ICASSP 2025},
year = {2025}
}