How Sampling Rate Affects Cross-Domain Transfer Learning for Video Description
Yu-Sheng Chou, Pai-Heng Hsiao, Shou-De Lin, Hong-Yuan Mark Liao
Abstract
Translating video to language is very challenging due to diversified video contents originated from multiple activities and complicated integration of spatio-temporal information. There are two urgent issues associated with the video-to-language translation problem. First, how to transfer knowledge learned from a more general dataset to a specific application domain dataset? Second, how to generate stable video captioning (or description) results under different sampling rates? In this paper, we propose a novel temporal embedding method to better retain temporal representation under different video sampling rates. We present a transfer learning method that combines a stacked LSTM encoder-decoder structure and a temporal embedding learning with soft-attention (TELSA) mechanism. We evaluate the proposed approach on two public datasets, including MSR-VTT and MSVD. The promising experimental results confirm the effectiveness of the proposed approach.
BibTeX
@inproceedings{icassp2018_howsamplingratea,
title = {How Sampling Rate Affects Cross-Domain Transfer Learning for Video Description},
author = {Yu-Sheng Chou and Pai-Heng Hsiao and Shou-De Lin and Hong-Yuan Mark Liao},
booktitle = {ICASSP 2018},
year = {2018}
}