ICASSP 2025accepted0 citations

Robust and Efficient Text-based Speech Editing using Noise Conditioning and Rectified Flow

Haowen Yin, Kai Wang, Hongli Yang, Hao Huang, Wushour Silamu

Abstract

Significant advancements have been made in text-based speech editing (TSE) for clear speech, but effectively editing the noise-contaminated speech remains a challenge. Background noise degrades the quality of generated speech, and edited speech that fails to maintain noise context consistency often sounds unnatural. We propose Reflow-TSE, a robust and efficient TSE model for noise-consistent speech editing. For noise robustness, 1) we design a noise condition module to extract frame-level noise sequences as the conditioning information, and 2) we introduce an enhanced context-conditioned prediction module that predicts masked noise sequences along with conventional duration and pitch, using these conditions to guide generation. For boosting efficiency, 3) we introduce the rectified flow model, leveraging speech context and predicted conditions to achieve high-quality editing with limited sampling steps. Experimental results show that with just two steps of sampling, Reflow-TSE achieves context-consistent noisy speech editing, a capability absent in other TSE models. Additionally, for clean speech, Reflow-TSE also matches or surpasses baseline models. Audio samples are available at https://hw-su.github.io.

BibTeX
@inproceedings{icassp2025_robustandefficie,
  title = {Robust and Efficient Text-based Speech Editing using Noise Conditioning and Rectified Flow},
  author = {Haowen Yin and Kai Wang and Hongli Yang and Hao Huang and Wushour Silamu},
  booktitle = {ICASSP 2025},
  year = {2025}
}