Face Landmark-based Speaker-independent Audio-visual Speech Enhancement in Multi-talker Environments
In this paper, we address the problem of enhancing the speech of a speaker of interest in a cocktail party scenario when visual information of the speaker of interest is available.Contrary to most previous studies, we do not learn visual features on the typically small audio-visual datasets, but use…