2025
Frame-Voyager: Learning to Query Frames for Video Large Language Models
ICLR 2025poster
Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame r…