ICASSP 2025accepted0 citations

Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition

Yufeng Yang, Hassan Taherian, Vahid Ahmadi Kalkhorani, DeLiang Wang

Abstract

Despite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing artifacts that degrade the ASR backend trained on clean speech. As a result, mainstream robust ASR systems train on noisy speech to avoid processing artifacts. In this work, we propose to decouple the training of the multi-channel speaker separation frontend and the ASR backend, with the latter trained only on clean speech. On SMS-WSJ, the proposed approach achieves a word error rate (WER) of 5.74%, outperforming the previous best by 14.3%. Furthermore, on recorded LibriCSS, we achieve the speaker-attributed WER of 3.86%, outperforming the previous best system trained on the same data by 24.8%. These state-of-the-art results suggest that decoupling speech separation and recognition is a potentially effective approach to robust ASR.

BibTeX
@inproceedings{icassp2025_elevatingrobusta,
  title = {Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition},
  author = {Yufeng Yang and Hassan Taherian and Vahid Ahmadi Kalkhorani and DeLiang Wang},
  booktitle = {ICASSP 2025},
  year = {2025}
}