Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone Arrays
Hongmei Guo, Linfeng Feng, Yijiang Chen, Xueqing Li, Boyu Zhu, Hao-Yu Wang, Xiao-Lei Zhang, Xuelong Li
Abstract
Speech separation using ad-hoc microphone arrays has been explored, but there is still significant room for improvement, especially in complex scenarios with varying channel conditions. Co-attention, a feature fusion mechanism, is widely used in multimodal fusion to capture the cooperation between modalities and enhance the representation of extracted features. In this paper, we propose a co-attention-based multi-channel model for speech separation with ad-hoc microphone arrays. The co-attention mechanism is integrated into the model to enhance the interaction between different speakers across multiple channels, enabling efficient channel fusion. To the best of our knowledge, this is the first work to apply co-attention for speech separation. Experimental results demonstrate that the proposed method significantly outperforms existing approaches, underscoring the importance of co-attention in optimizing channel fusion for speech separation in challenging acoustic environments.
BibTeX
@inproceedings{icassp2025_coattentionbased,
title = {Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone Arrays},
author = {Hongmei Guo and Linfeng Feng and Yijiang Chen and Xueqing Li and Boyu Zhu and Hao-Yu Wang and Xiao-Lei Zhang and Xuelong Li},
booktitle = {ICASSP 2025},
year = {2025}
}