2018
Watch, Listen Once, and Sync: Audio-Visual Synchronization With Multi-Modal Regression Cnn
ICASSP 2018accepted
Recovering audio-visual synchronization is an important task in the field of visual speech processing. In this paper, we present a multi-modal regression model that uses a convolutional neural network (CNN) for recovering audio-visual synchronization of single-person speech videos. The proposed mode…