2026
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
CVPR 2026
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world