2025
ReWind: Understanding Long Videos with Instructed Learnable Memory
CVPR 2025poster
Vision-Language Models (VLMs) are crucial for real-world applications that require understanding textual and visual information. However, existing VLMs face multiple challenges in processing long videos, including computational inefficiency, memory limitations, and difficulties maintaining coherent…