2022
VoViT: Low Latency Graph-Based Audio-Visual Voice Separation Transformer
ECCV 2022poster
"This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with a lightweight graph convolutional network that processes fa…