Personalized Speech Enhancement without User Enrollment for Real-World Audio Replay Scenarios
Haoran Wei, Shilin Wang, Yanhua Long
Abstract
Many speech enhancement (SE) approaches have been proposed to deal with cocktail party problem. Personalized speech enhancement (PSE) approaches improve SE performance by utilizing user enrollment speech. However, PSE requires users to record additional clean audio for registration, which can be redundant works or impractical for many real-world scenarios. For instance, in personal devices and Vloggers’ audio playback scenarios, there are already many video/audios available recorded under different noise types and SNR conditions, but without any target speaker pre-registered speech. To better utilize information from existing video/audio stock, this paper propose a novel speech enhancement approach that integrates PSE methods without requiring pre-registered user speech. With user adaptation and noise adaptation training modules, the proposed approach automatically selects high-quality speech segments to assist in denoising low-quality speech segments. Additionally, two test sets were collected to evaluate the performance in the aforementioned scenarios. Experimental results demonstrate that the proposed approach outperforms the corresponding SE methods in both objective and subjective evaluation metrics.
BibTeX
@inproceedings{icassp2025_personalizedspee,
title = {Personalized Speech Enhancement without User Enrollment for Real-World Audio Replay Scenarios},
author = {Haoran Wei and Shilin Wang and Yanhua Long},
booktitle = {ICASSP 2025},
year = {2025}
}