ICASSP 2026poster0 citations

RELUNET: RELATIVE CHANNEL FUSION U-NET FOR MULTICHANNEL SPEECH ENHANCEMENT

Ibrahim Aldarmaki, Bhiskha Raj, Hanan Aldarmaki

Abstract

Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate the channels during later stages of the network. In this paper, we propose a novel modification of these models by incorporating relative information from the outset, where each channel is processed in conjunction with a reference channel through stacking. This input strategy exploits comparative differences to adaptively fuse information between channels, thereby capturing crucial spatial information and enhancing the overall performance. The experiments conducted on the CHiME-3 dataset demonstrate improvements in speech enhancement metrics across various architectures.

BibTeX
@inproceedings{icassp2026_relunetrelativec,
  title = {RELUNET: RELATIVE CHANNEL FUSION U-NET FOR MULTICHANNEL SPEECH ENHANCEMENT},
  author = {Ibrahim Aldarmaki and Bhiskha Raj and Hanan Aldarmaki},
  booktitle = {ICASSP 2026},
  year = {2026}
}