2026
Efficient Frame Selection for Long Video Understanding via Reinforcement Learning
CVPR 2026
Recent advances in Multimodal Large Language Models (MLLMs) have led to significant progress in video understanding. Due to limited context windows and computational overhead, most MLLMs adopt uniform frame sampling. This approach is at high risk of missing critical visual information and constrains