ICASSP 2024accepted0 citations

Lightweight Multi-Axial Transformer with Frequency Prompt for Single Channel Speech Enhancement

Xingwei Liang, Zehua Zhang, Mingjiang Wang, Ruifeng Xu

Abstract

Time-frequency analysis in single-channel speech enhancement has received considerable attention. While Transformer-based architectures are gaining traction, their computational burden can be substantial, especially when dealing with longer speech samples. To address this, our research introduces the lightweight multi-axial Transformer (LMA-Transformer) optimized for low computational overhead while efficiently extracting features along both temporal and frequency axes. Major to our approach is the Temporal/Frequency MultiDConv head self-attention module (T/F-MDHSA), which not only reduces computational costs but also improves the Transformer’s capability to utilize local features effectively. Moreover, we introduce the frequency prompt block, designed to dynamically guide the recovery of frequency features in speech signals that have experienced varying levels of degradation. Compared to state-of-the-art models, our model has competitive performance with 3.40 PESQ, 95.8% STOI, and 10.15 SSNR on the VoiceBank + Demand dataset.

BibTeX
@inproceedings{icassp2024_lightweightmulti,
  title = {Lightweight Multi-Axial Transformer with Frequency Prompt for Single Channel Speech Enhancement},
  author = {Xingwei Liang and Zehua Zhang and Mingjiang Wang and Ruifeng Xu},
  booktitle = {ICASSP 2024},
  year = {2024}
}
Lightweight Multi-Axial Transformer with Frequency Prompt for Single Channel Speech Enhancement · ICASSP 2024