2025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
NeurIPS 2025poster
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability…