ICASSP 2025accepted0 citations

Modifying Flow Matching for Generative Speech Enhancement

Roman Korostik, Rauf Nasretdinov, Ante Jukic

Abstract

Diffusion-based generative models have been shown to be highly effective in various speech enhancement tasks. This work presents an analysis of a flow matching-based framework for generative speech enhancement as a simpler alternative to diffusion. Four different modifications to flow matching are proposed, employing an informed prior, a data prediction loss, deterministic inference, and early stopping. The proposed variants are evaluated on speech denoising, demonstrating performance comparable to a previous state-of-the art model using the same data setup. Through ablation studies, an efficient deterministic one-step inference configuration is proposed, which does not require any advanced training techniques such as pre-training or distillation. The proposed variants are also evaluated on speech dereverberation, demonstrating that stochastic inference without informed prior is preferable for this task.

BibTeX
@inproceedings{icassp2025_modifyingflowmat,
  title = {Modifying Flow Matching for Generative Speech Enhancement},
  author = {Roman Korostik and Rauf Nasretdinov and Ante Jukic},
  booktitle = {ICASSP 2025},
  year = {2025}
}