2026
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
AAAI 2026technical
Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and resource constraints during both training and inference. Alth