2026
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
CVPR 2026
Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically perceive temporal dynamics via positional encoding, text-based timestamps, or visu