2024
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
ECCV 2024poster
"In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent representation learned from a pretrained generative T2V model encapsulates rich semantics and coherent temporal correspo…