2026
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
CVPR 2026
Long video understanding is a key challenge that plagues the advancement of Multimodal Large language Models (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and proposed a novel and training-free approach, termed Flexible Memory (FlexMem). In principle,