ICASSP 2018accepted0 citations
Watch, Listen Once, and Sync: Audio-Visual Synchronization With Multi-Modal Regression Cnn
Abstract
Recovering audio-visual synchronization is an important task in the field of visual speech processing. In this paper, we present a multi-modal regression model that uses a convolutional neural network (CNN) for recovering audio-visual synchronization of single-person speech videos. The proposed model takes audio and visual features of multiple frames as the input and predicts a drifted frame number of the audiovisual pair which we input. We treat this synchronization task as a regression problem. Thus, the model does not need to search with a sliding window which would increase the computational cost. Experimental results show that the proposed method outperforms other baseline methods for recovered accuracy and computational cost.
BibTeX
@inproceedings{icassp2018_watchlistenoncea,
title = {Watch, Listen Once, and Sync: Audio-Visual Synchronization With Multi-Modal Regression Cnn},
author = {Toshiki Kikuchi and Yuko Ozasa},
booktitle = {ICASSP 2018},
year = {2018}
}