Causal Speech Enhancement Based on a Two-Branch Nested U-Net Architecture Using Self-Supervised Speech Embeddings
This paper presents a causal speech enhancement (SE) model based on a complex two-branch nested U-Net architecture (CNUNet-TB) combined with a two-stage (TS) training method that leverages speech embeddings from a large self-supervised speech representation learning (SRL) model. The proposed archite…