Prelude echoes Finale: Video Domain Adaptation with Fine-grained Temporal Consistency
Shuxian Wang, Mengmeng Jing, Xianlong Tian, Yuguo Hu, Lin Zuo
Abstract
The prelude and finale of one video usually share similar themes and content, with the finale often echoing or deepening the concepts introduced in the prelude. This temporal correlation enhances the understanding of video content. Existing video domain adaptation methods, however, ignore this fine-grained temporal structure in the clip level, focusing instead on coarser temporal consistency at the video level. To make the prelude echo the finale, we propose to optimize fine-grained temporal consistency for video domain adaptation. Specifically, we first extract the prelude (the first n clips) and the finale (the last n clips) from the same video. Then, we maximize the feature similarity of the prelude and the finale, which enhances the connections between clips and achieves the temporal consistency of each video across source and target domains. On the other hand, we generate the cross-domain fusion features and optimize their discriminability to make the source domain features echo the target. As a result, the source and target domains are aligned. Extensive experiments on video domain adaptation benchmarks demonstrate the effectiveness of our method.
BibTeX
@inproceedings{icassp2025_preludeechoesfin,
title = {Prelude echoes Finale: Video Domain Adaptation with Fine-grained Temporal Consistency},
author = {Shuxian Wang and Mengmeng Jing and Xianlong Tian and Yuguo Hu and Lin Zuo},
booktitle = {ICASSP 2025},
year = {2025}
}