2025
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
NeurIPS 2025poster
Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-LLM) architectures have critical limitations in temporal understanding, strugglin…