2023
Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding
CVPR 2023poster
Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the la…