2024
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
ICLR 2024poster
Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA) models operate in the time domain. However, their overly sim…