ICASSP 2025accepted0 citations

Speech Enhancement with MAP-based Training for Robust ASR

You-Jin Li, Rong Chao, Borching Su, Yu Tsao

Abstract

To improve noise robustness in automatic speech recognition (ASR), a common strategy is to employ speech enhancement (SE) models as front-ends for ASR systems. However, SE models often introduce artifacts into enhanced signals, which can degrade ASR performance, particularly when the SE and ASR models are trained separately. Although various methods have been proposed to address this issue, they often come at the expense of increasing implementation complexity. Hence, this study proposes a maximum a posteriori (MAP) algorithm for training SE models by incorporating the posterior probability of clean speech, given the enhanced speech, into the loss function. Experimental results show that our method enhances the compatibility between SE and ASR models in both in-domain and out-of-domain testing scenarios, notably improving ASR performance. The proposed method does not require prior knowledge of ASR models or speech content during training or inference, nor does it involve additional post-processing steps.

BibTeX
@inproceedings{icassp2025_speechenhancemen,
  title = {Speech Enhancement with MAP-based Training for Robust ASR},
  author = {You-Jin Li and Rong Chao and Borching Su and Yu Tsao},
  booktitle = {ICASSP 2025},
  year = {2025}
}