SepMamba: State-Space Models for Speaker Separation Using Mamba
Thor Højhus Avenstrup, Boldizsár Elek, István László Mádi, András Bence Schin, Morten Mørup, Bjørn Sand Jensen, Kenny Falkær Olsen
Abstract
Deep learning-based single-channel speaker separation has improved significantly in recent years in large part due to the introduction of the transformer-based attention mechanism. However, these improvements come with intense computational demands, precluding their use in many practical applications. As a computationally efficient alternative with similar modeling capabilities, Mamba was recently introduced. We propose Sep-Mamba, a U-Net-based architecture composed of bidirectional Mamba layers. We find that our approach outperforms similarly-sized prominent models — including transformer-based models — on the WSJ0 2-speaker dataset while enjoying significant computational benefits in terms of multiply-accumulates, peak memory usage, and wall-clock time. We additionally report strong results for causal variants of SepMamba. Our approach provides a computationally favorable alternative to transformer-based architectures for deep speech separation.
BibTeX
@inproceedings{icassp2025_sepmambastatespa,
title = {SepMamba: State-Space Models for Speaker Separation Using Mamba},
author = {Thor Højhus Avenstrup and Boldizsár Elek and István László Mádi and András Bence Schin and Morten Mørup and Bjørn Sand Jensen and Kenny Falkær Olsen},
booktitle = {ICASSP 2025},
year = {2025}
}