ICASSP 2022accepted0 citations
Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement
Lianwu Chen, Chenglin Xu, Xu Zhang, Xinlei Ren, Xiguang Zheng, Chen Zhang, Liang Guo, Bing Yu
Abstract
Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, personalized speech enhancement. As shown in the evaluation results, this is achieved by employing a multi-stage and multi-loss training architecture that incorporates the recently proposed two-step structure, ASR loss produced by a back-end ASR encoder, and the speaker extraction network.
BibTeX
@inproceedings{icassp2022_multistageandmul,
title = {Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement},
author = {Lianwu Chen and Chenglin Xu and Xu Zhang and Xinlei Ren and Xiguang Zheng and Chen Zhang and Liang Guo and Bing Yu},
booktitle = {ICASSP 2022},
year = {2022}
}