2026
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
CVPR 2026
In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented approaches are often used to process long videos, they usually compress each frame independently and therefore fail to