ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning
Video understanding is essential for multimodal large language models (MLLMs) to interact effectively with users and the real world. However, analyzing long videos remains a major challenge due to the lack of high-quality video instruction data and effective training strategies. In this paper, we in