FLOWSE-GRPO: TRAINING FLOW MATCHING SPEECH ENHANCEMENT VIA ONLINE REINFORCEMENT LEARNING
Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-training alignment via reinforcement learning (RL) effectively aligns generative models with human preferences and downstr…