2025
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
CVPR 2025poster
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding,…