ICASSP 2022accepted0 citations

Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement

Lianwu Chen, Chenglin Xu, Xu Zhang, Xinlei Ren, Xiguang Zheng, Chen Zhang, Liang Guo, Bing Yu

Abstract

Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, personalized speech enhancement. As shown in the evaluation results, this is achieved by employing a multi-stage and multi-loss training architecture that incorporates the recently proposed two-step structure, ASR loss produced by a back-end ASR encoder, and the speaker extraction network.

BibTeX
@inproceedings{icassp2022_multistageandmul,
  title = {Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement},
  author = {Lianwu Chen and Chenglin Xu and Xu Zhang and Xinlei Ren and Xiguang Zheng and Chen Zhang and Liang Guo and Bing Yu},
  booktitle = {ICASSP 2022},
  year = {2022}
}