2024
Weakly-Supervised Spatio-Temporal Video Grounding with Variational Cross-Modal Alignment
ECCV 2024poster
"This paper explores the spatio-temporal video grounding (STVG) task, which aims at localizing a particular object corresponding to a given textual description in an untrimmed video. Existing approaches mainly resort to object-level manual annotations as the supervision for addressing this challengi…