Advancing Non-intrusive Suppression on Enhancement Distortion for Noise Robust ASR
Wei Wang, Siyi Zhao, Yanmin Qian
Abstract
Recent advancements in speech enhancement (SE) techniques have greatly improved speech clarity and intelligibility in challenging acoustic environments. However, integrating SE into automatic speech recognition (ASR) systems often results in performance degradation due to artifacts introduced during the enhancement process. While various methods have enhanced recognition accuracy in SE-ASR systems, they often require fine-tuning or re-training of SE or ASR models, which is impractical in many real-world applications. In this paper, we propose a lightweight distortion suppression (DS) network that addresses these artifacts without modifying the SE or ASR models, treating them as fixed black boxes. The DS module operates on the time-frequency (T-F) bands of the original and enhanced complex spectrograms, efficiently compensating for SE distortions using the original T-F information. We validate our approach through experiments on both Mandarin and English ASR tasks using monaural and multi-channel SE frontends, across various ASR backends. Results show that the DS module significantly improves the performance of SE-ASR systems, even when used with robust commercial ASR backends.
BibTeX
@inproceedings{icassp2025_advancingnonintr,
title = {Advancing Non-intrusive Suppression on Enhancement Distortion for Noise Robust ASR},
author = {Wei Wang and Siyi Zhao and Yanmin Qian},
booktitle = {ICASSP 2025},
year = {2025}
}