ICASSP 2023accepted0 citations

A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement

Haitao Xu, Liangfa Wei, Jie Zhang, Jianming Yang, Yannan Wang, Tian Gao, Xin Fang, Li-Rong Dai

Abstract

Audio-visual speech enhancement (AVSE) was shown to be superior over conventional audio-only counterpart for improving the speech quality. However, most existing AVSE models are heavyweight in the sense of parameter count, which is inappropriate for the deployment and practical applications. In this paper, we therefore present a lightweight AVSE approach (called M3Net) by incorporating several multi-modality, multi-scale and multi-branch strategies. Three multi-scale techniques are designed for the visual and audio streams, including multi-scale average pooling (MSAP), multi-scale ResNet (MSResNet) and multi-scale short time Fourier transform (MSSTFT). It is shown that each multi-scale module positively contributes to the performance. Also, we consider four skip connections for the audio-visual feature aggregation, which have a great complementary effect on the designed multi-scale techniques. Experimental results show that these techniques are flexible in combination with existing approaches, and more importantly obtain a comparable performance with a smaller model size compared to the heavyweight networks.

BibTeX
@inproceedings{icassp2023_amultiscalefeatu,
  title = {A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement},
  author = {Haitao Xu and Liangfa Wei and Jie Zhang and Jianming Yang and Yannan Wang and Tian Gao and Xin Fang and Li-Rong Dai},
  booktitle = {ICASSP 2023},
  year = {2023}
}